Friday, August 7, 2009

Finding the Next Dot-Com

Its seems like everyone and their mom has a domain name nowadays, and domain squatters have the rest.

The land grab is speeding up. VeriSign says that more than 2.4m new .coms are registered every month. Virtual real estate is just a name, but domain auctions sometimes make actual real estate look cheap--and even if you aren't paying $7.4m for "business.com", a lot of domain names are expensive. If it's pronounceable, paying a few thousand bucks is apparently normal.

All this squatting and grabbing makes life harder for people who actually want to build a web presence. Just look at the names of some of the newer companies around here: Ooyala, anyone? DealKat? SaaSure? That first one reads like a misspelling of a French exclamation. The second is a misspelling. The third is a pun on SaaS, which is usually pronounced "sass", so it sounds like a lisp ("sass-sure"?).

These guys seem to have gotten what lots of people want: an unregistered domain name. Those sell for about $10, renewed yearly; the details depend on the specific TLD you want. In any case, they're vastly cheaper than domains that are already registered, whether by previous legitimate users or by squatters.

After spending half an hour recently typing stuff into whois, looking for the coveted "not found" that would tell me a name was unregistered, I found nothing. I thought: there's got to be a better way...

So I wrote a few quick-and-dirty programs that generate domain names, then used a shell script to test them automatically and pick out the unregistered ones. I wanted domains that were reasonably short and looked and sounded like English words. Since all the tricks I tried to that end rely on a large English word list, and since I didn't feel like waiting very much, I wrote these mini-programs in C++.



Attempt #1 goes through a list of English words and builds a map of all each 'syllable' along with how often that syllable occurs. For syllables, I just used all the substrings in each word of 2-5 letters that had at least one vowel. The most common were es, in, er, ed, ing...

Then, I combined two syllables at a time, starting with the the most common and the second most common: esin.com, ines.com, eser.com, eres.com...

The names got a bit more complicated after that, but they weren't very wordlike. Worse, the first 1000 were all four or five letters long, so despite sounding like gibberish, almost all of them were taken.

Attempt #1 results

500 domains
53 of which happen to be in the original wordlist. None of those was available.
Of the remaining 447, 2 were available: 0.4%

I needed something better...



Attempt #2 is a Markov model of the words in the dictionary. After a bit of experimenting, I decided to use a Markov model. For every three-char sequence in the input word list, I kept track of all the characters that come next. I included the null character at the end of each word, to model the length of English words as well as the characters they're made of.

To generate names, I sampled the most common three-character sequences at the beginning of dictionary words. For each sequence, I randomly picked from the characters that could come next, and repeated that until I got to a null character.

This one worked a lot better:

dislew.com
comm.com
trati.com
fore.com
staying.com
impo.com
recre.com
stri.com
gravel.com
miss.com
...

I especially like comm.com. Maybe a communications or PR agency? It's a sweet domain.

These names were generally a bit longer, but pronounceable. A lot more of them were available--it seems like raw length is what the squatters really care about.

Attempt #2 results:

500 domains
122 of which happen to be in the original wordlist. None of those were available.
Of the remaining 378, 61 were available: 16.4%

Possibilities abound. I could..
  • Artificially crank up the probability of the null character for every seed in the Markov model, to make it produce shorter domain names while still trying to keep them wordlike.
  • Exclude long words from the wordlist and then build the Markov model, also in an effort to produce shorter domain names.
  • Use the Markov model to generate two words at a time, and simply concatenate them, also to get more compound-word domain names.
  • Repeat the experiment for languages other than English.
  • Run a bunch of popular domains and less popular ones through the Markov model to see how they score. Are domains that look like words more popular than those that don't?
  • Compare some sites with a lot of direct traffic (people typing the domain into their URL bars, instead of going through a link or a search engine) and see if those tend to be more wordlike than their less typeable counterparts.

I might try some of this stuff. If I do, I'll keep you posted. Peace...

Sunday, April 19, 2009

I'll be BAWK



Last weekend, we went on an epic roadtrip through California.

It's an annual section leader tradition, and this was its fourth year. On Friday afternoon, we piled into cars with food, sleeping bags, laptops, and cameras, and started our trip. The goal? Solving puzzles and getting places as fast as possible. BAWK is framed as an adventure game. The theme? Always a B-movie. This year it was Cannibal Women of the Avocado Jungle of Death The route? TBD. We go nowhere fast, then hang out and have fun.



It turned out that we were driving from Stanford to within 90 miles of Mexico on that first afternoon. It was a beautiful drive through the Sacramento Valley and the Mohave Desert.

Stop 1: the Salton Sea.



A rainbow on the way to Dinosaur Point (San Luis Resovoir).

We camped on the Salton Sea our first night. In the morning, we went to an abandoned trailer park that had sunk into a wet marsh. Pictures coming soon.



Antelope Valley California Poppy Reserve.



Joshua Tree National Park.

On the way back, we made it to Big Sur before sunset. It was gorgeous.



Brandon, the B in BAWK and roadtrip planning genius.



The sun.




The sea.





The beach. We swam here.

I love the desert. I love the backroads, the cogs under the surface of society, the small towns that grow our food, transport our goods, generate our electricity. I love the freedom of driving with friends through the middle of nowhere, going wherever we want to go.



California is beautiful.

Tuesday, January 6, 2009

Winning science fairs

... I've never done it, but I did get a "third award" at Intel ISEF '07. That's International Science + Engineering Fair, and it was one of the most fun things I've ever done. My friend Tom and I built the project, took it to regional fairs, improved it between each fair, and of course partied the nerdy way for a week at the main event. Our project was on Stirling Engines, building four from scratch. The last two worked.

We built the engines using parts from my garage and from Home Depot. The last engine also included two borosilicate glass-graphite piston cylinders generously donated by Airpot. The Riverton Science Department hooked us up with real-time pressure and rotation sensors. The end result was an engine that could run on a candle and a faster one that ran off of a propane stove; we measured the thermodynamic cycle of the second engine, creating and experimental PV diagram.



My partner-in-crime Sawyer just facebooked:
Hey! You went to ISEF a while ago, didn't you? Any advice on choosing a project?


Sure! In my unhumble opinion, engineering is the way to go.
Pick something fun or cool, because then you'll spend time on it and make it great. Also, pick something that you can build. I have way better experience with that than doing pure science or math, at least in the context of a science fair. (Science fairs are really science/math/engineering fairs; I like the engineering part best.) If you do pure math or science, then you'll need a mentor. Many entrants in those categories, even at Utah-level fairs, partner with university professors; needless to say, it's very hard to compete if you're going it alone. Building stuff is different; you need a lot less arcane knowledge. I wrote some of the programs I'm most proud of before I had any formal CS training; my friend Nathan built a car -- from scratch -- that does 0-60 in under four seconds -- before he went to Stanford and took introductory mechanical engineering. If you're excited about it, if you persevere, and if you're willing to use Wikipedia a lot, you can learn as you go and build really cool stuff.

I also recommend building something physical. You can enter a computer science project, but you'll have plenty of chances to code. In my experience, the clubs, groups, and other orgs you'll be part of in college never have enough programmers. Web developers are especially in demand -- you'll spend lots of time sitting in front of a screen hacking apps into existence. Science fairs are your chance to design, machine, tinker with, and generally build stuff, and get recognition for it. Companies sponsor you. Teachers love it and are often very helpful. It's a very good deal, and one you really don't get once you're out of school. I can't overstate how great an experience my project was. Find something interesting and build it. You'll rock the science fairs and get more out of it than you can imagine.


(Tom and I at Sandia National Labs. One of our judges invited us to come see his project there: harvesting solar energy with Stirling engines. The dish focuses sunlight onto the hot end of a 25KW engine. The other side is cooled with a radiator.)

Sunday, December 21, 2008

Dr. Markov

Last week, in CS106X, I programmed Boggle.

Boggle is played with a set of cubes that have letters on each face. You shake them up, put them in a grid, and try to find words made of adjacent dice faster than the other players can.

Computer Boggle is different. It gives you time, letting you find as many words as you can, and then beats your score by an order of magnitude or so by finding all the remaining words on the board that are in the (huge) Scrabble player's dictionary. You can also play 5x5 "BigBoggle" instead of the standard 4x4, in which case the computer generally beats you even more thoroughly.

One of the cool things about CS106 is that they don't just want solutions, though. It's good style to extend them in some way, and solve some additional problems that weren't in the assignment. For Boggle, I wanted to write code that would find boards with unusually few (or many) words, using only the standard cubes.

So I tried Markov chain Monte Carlo, which I'd Wikied a long, long time ago when I was really into raytracing.

In a nutshell: the algorithm tries to find a global minimum of a function that can be sampled and which is roughly continuous, but which may be in many dimensions and may not have any other nice properties.

You start with a guess.
At each step, you find a distribution centered at the current guess and randomly sample it. If you know what the global minimum should be (e.g. zero) the variance of the distribution will vary with the cost (the value of the function at the current guess), so that if you're far away from the minimum, you're sampling over a wide area, while if you're getting close, you make more incremental guesses. Each 'probe' sampled this way is evaluated, and then you randomly decide whether to accept or reject it, depending on its cost. If the probe cost is significantly lower than the current cost, then you almost certainly accept it, and if the cost is significantly higher, then you almost always reject; in between, it's a toss-up: if you accept, then that probe becomes your current guess, and if you reject, then you keep probing.

Doing all this with Boggle boards took a bit of approximation and rule-of-thumb coding because I don't know a way to actually define a probability distribution over the set of possible Boggle boards. Instead, I defined 'swapping' as taking two cubes at random, switching their places, and randomly rotating each one. I started with a standard, randomly-generated board and, after each evaluation, did a number of swaps proportional to my cost function, with a maximum of about ten. When the guesses started getting so good that I wasn't switching any cubes, I picked single cubes and rotated them randomly to create new probes. The idea is similar: the closer you get to a global minimum, the closer your probes are to your current guess.

I made two cost functions, one preferring boards with lots of words and and the other preferring boards with few words. I already had a function that could quickly search a Boggle board and find all the words on it.

The result: it worked. I was surprised.

BigBoggle boards tend to have around 150 words on them. Here's one with no words at all:


On regular boggle boards, the computer will score ~90. Here, it scored 917:


The code is private because of academic policy, but if you want to play with the program, it's here.

Markov chain Monte Carlo rules.

Thursday, December 4, 2008

Timesinks

The internet makes lots of tasks infinitely easier than they used to be. Want an executive summary of the War of 1812? Check Wikipedia. Want to know more? Keep reading. If you're really hardcore, the article lists eighty-two scholarly references. Need an obscure book printed by some shop in Khazakstan? Amazon can probably ship it to you for free in ten days or so. Want to convert a kilowatt hour into BTUs? Google "a kilowatt hour in btus". I have no experience here, but I bet that before the web, lots of these tasks would have taken minutes or hours rather than seconds.

Isn't it strange, though, that such a powerful tool for getting stuff done gets used so extensively for the purpose of not getting stuff done? I can't think of another way to explain the fifty-or-so bits of chatlist spam that land in my inbox every day, the SuperPokes, the pirates, the ninjas, the lolcats, and everything else along those lines. Even respectable, otherwise interesting sites like Digg and Youtube are filled with mildly funny, mildly shocking, and mildly sexual banalities.
A while ago, I posted about Twitter, the flagship time-waster-in-chief of all web 2.0 apps. Never before has humanity been this effective at wasting time.

I don't think that lost time is the real problem, though. When the novelty wears off and everything else they could have done with all that time sinks in, people will stop. I'm worried about money.

Parallel to this Cambrian explosion of "web 2.0" apps, replete with tweets, pokes, nudges, posts, threads, and lists, there have been all sorts of new revenue models, further and further removed from the sale of actual goods and services to consumers. Take Slide, Inc, which recently got a $500m valuation from some venture capitalists. It makes slideshow widgets (and little gems like "SuperPoke Pets!") which it plans to monetize by integrating targeted ads into users' slideshows. Many of the ads I see (and ignore) all the time are already advertise other ad-supported sites that don't actually sell anything. It's only a matter of time before I see some ads for Slide.

I think that the system by which dotcoms are funded today looks suspiciously like the way they were funded in the nineties. People are spending way too much money building castles in the air.

Someday, people will tire of spending so much time connected to each other, wasting each others time. The torrent of tweets, pokes, FWD:fwd:Fwd:Fwd:funny pics, and all their unholy brethren will start slowing down as people take back the time that used to be theirs. Maybe someday, sending pointless emails to a chatlist will have the same social stigma that shouting across a crowded room might have now. Maybe the parents of 2015 will tell their kids to finish their homework before they comment on each other's blogs. Maybe they'll warn their kids that too much Twitter rots your brain, or that Facebook bullies are probably just looking for attention and are best left ignored. Maybe someday, I'll log into Facebook and nobody will care what my hooker name would be. I can't wait for that day to come!

Let me skip back for a moment, to the first stock I ever owned ~ two shares of Microsoft, back in 1998 when I was in second grade. My best friend's dad thought that a market in which grade-schoolers own stock couldn't possibly be sustainable. He was right. Today, we've got a market in which investors bet billions on web startups with Rube Goldberg business plans, with almost no connection to reality. They're betting on people barely older than myself trying to capitalize on the time-wasting habits of people, often a lot younger than me, by showing them ads. As people become savvier and the web becomes fully integrated into our culture, my guess is that many dotcoms will be winnowed out in a second crash.

Thursday, November 6, 2008

President Obama


This past weekend has been a blur. Some of my fellow students and I went to Las Vegas to help swing Nevada to Obama's side. We drove all day on Saturday; we worked on Sunday and Monday. I woke up at five on Election Day to put up door hangers, stayed up until one in the morning with the most energetic crowd I've ever seen, and woke up at five yesterday to fly back to California.

All the campaigning and partying, while fun, was on some level superficial. I got into some heated arguments with a junior in our group who seemed to be channeling Karl Marx, and into some more interesting arguments with other, more moderate fellow campaigners. For the most part, though, we focused on the very vague messages of "hope" and "change", and passed out policy fliers to the handful of people who wanted to know more. Nobody really asked the question of what America would become (or, now, will become) under an Obama presidency.

I've been looking for some good opinions. I thought that this article had an interesting take on it.

Like lots of what's been published since Election Day, it focuses pretty heavily on race. I'm more interested in what will happen to the economy. At least for the national debt, change seems like a good thing. Obama's tax plan has always bothered me, though.

Like most of us, I'll wait and see.

If there's one thing thats unambiguously good, it's that people care now. Lots of them who never voted before, voted this time. Students stopped ignoring politics. If you'd asked me just a few weeks ago, during New Student Orientation, what I thought about the next four years, I would have been gung-ho and pretty self-centered: Yes, Stanford will be fantastic. I'm still excited about the next four years, but like lots of other people I think about them in a broader way. What are we doing to help bring change?
 
Views