In the stone age, primitive humans wrote on clay tablets and carried memorized poems throughout generations. In the computer age, we, perhaps still primitive, write on laptops. To synthesize history and the future, we’re starting to talk to our computers, who will record our words and meanings forevermore.
Our computers are attaining new capabilities every minute. With the invention of the transformer, a neural network that learns by predicting what comes next in a sequence, the architecture behind systems like ChatGPT, our machines are gaining new capabilities: talking in natural language, reasoning beyond basic arithmetic, image generation. These skills are not just impacting what computers can do. Media theorist Marshall McLuhan, who spent his career arguing that novel technologies reshaped the people who used them, put it, “We shape our tools and thereafter our tools shape us.” This change happens along two axes: capabilities and interfaces. An interface is the layer through which a person interacts with a computer, issuing instructions and receiving information in return. The fact that new computer capabilities reshape us is taken for granted: search engines redrew the line between what we remember and what we look up; GPS made navigating through cities easier at the cost of learning the geography for ourselves. Interfaces like the keyboard, the mouse, the trackpad, and the desktop have defined the history of computing, and now voice – the most human of communication systems – will go from an accessibility tool, to one of the primary ways we communicate with computers.
When we think of using a computer, we usually think of dragging a mouse and pointing it at something on a screen, or typing on a keyboard and hitting ‘enter.’ Look down at your keyboard, a descendant of the typewriter, and look at the letter Q: the keys to the right of it spell out “QWERTY.” The “QWERTY” keyboard, as it came to be called, was invented by Christopher Latham Sholes, a newspaperman in Milwaukee who, through roughly fifty prototypes spanning from 1867 to 1873, helped build one of the first practical typewriters. Typewriters used to jam when neighboring keys were hit in quick succession, so Sholes separated the most common letter combinations — and QWERTY was born.
The punched card predates the keyboard by decades. Jacquard looms (named after their inventor, Joseph Marie Jacquard) used punch cards to store weaving patterns — each card encoded a pattern for the machine to weave, and changing the card changed the pattern of the textile. But it wasn’t until Herman Hollerith adapted the idea for the 1890 US census that cards became a way to feed data to a machine, each hole corresponding to a demographic category. It wasn’t until digital computers had a command line — a plain window with a blinking cursor where you type instructions and the computer prints back the results, as opposed to the GUI (graphical user interface) that lets you click around icons onscreen — that information exchange between man and machine had a tight feedback loop. But programmers still had to memorize text commands. In the late 1960s, the command DWIM (Do What I Mean) was introduced in the BBN Lisp programming environment as a way to correct mistyped commands rather than just erroring out. The idea later lived on in programmer culture through the text editor Emacs with commands like `comment-dwim,` where a single command adapts to what the user seems to intend, rather than stumbling through errors.
In December 1968, at Brooks Hall — an event space beneath Civic Center Plaza in San Francisco — one of the most important events in the history of computers took place, a 90-minute presentation retroactively titled “The Mother of All Demos.” Douglas Engelbart, a researcher at the Stanford Research Institute (SRI), introduced the world to the field of human-computer interaction and the fundamentals of what would become personal computing. The demo was the first public showing of windowing, hypertext, graphics, command input, video conferencing, the word processor, dynamic file linking, version control, and, most famous of all, the mouse. (Although Engelbart invented few of these himself; his achievement was fusing them into one live, interactive system.) Before the mouse, using a computer often meant working entirely through a keyboard and memorizing text commands. The mouse brought to computing the uniquely human experience of pointing — one of our earliest modes of communication. It got the name because the early device, with a cord trailing out the back, reminded Engelbart’s team of a mouse with a long tail.
Unlike the artificial arrangement of QWERTY, which the human had to adapt to, the mouse set the trend toward a more natural way of interacting with computers. The ape that once could only point was now pointing to move text, arrange windows, write programs and build.
The mouse let us point at a screen. The next leap, that of Steve Jobs, let us touch it. In 2007, he introduced the iPhone to the world: a phone with internet access and a set of “apps” you interacted with by touch. It was unlike other phones in that it had no physical keyboard — Jobs’ insight was that sometimes, you don’t want the keyboard, you want as much screen as your pocket can fit. The iPhone’s capacitive multi-touch screen replaced the stylus with the best pointing device we have: our finger. With a natural grammar — pinch to zoom, flick to scroll with momentum, lists that bounce at their edges as if the software had mass — Apple extended the human body’s instincts to the machine.
Underneath the glass of the first iPhone, the software was doing the work. The iPhone operating system (now called iOS but originally labeled as a mobile version of Apple’s desktop OS X called “OS X”) made the experience magical: touching a photo was like touching the photo itself. It was the purest expression of Jobs’ design creed — that using a computer should be so intuitive it needed no manual.
The iPhone’s victory looks obvious in retrospect, but for years it was not clear at all — phones kept iterating with mechanical keyboards that would flip, rotate, or otherwise stay fixed. What’s safe to say is that every smartphone since has been built in its image.
But touch only solved half of Jobs’ creed. A screen you point at is intuitive, but humans don’t communicate through touch alone. Humans most naturally communicate by speaking. In 2011, Apple introduced Siri, spun out of SRI, formerly the Stanford Research Institute, with its iPhone 4S. Siri was one of the first voice assistants most people had ever used, built on early natural language processing (NLP) and machine learning research. Revolutionary as it was — it was, after all, the first time many customers could talk to a machine and get an answer — talking to Siri was far from smooth. The assistant matched what you said against a fixed set of commands, like setting a timer or checking the weather. Outside of that, it directed users to web search. Its few agentic capabilities remain largely confined to select Apple apps, unable to (or unwilling to) translate across the digital ecosystem. The naturalness of communication wasn’t there because the AI wasn’t either.