The Unknown History of Excel and AI
You don't know Clippy.
Memoirs of the once and never cartoon king of AI.
The road to Clippy improbably started in the most serious of applications, Excel 5.0.
In 1993, we had a problem: even Excel 4.0 had gotten so feature-rich that finding anything meant reading manuals or hunting down a product guru. We wanted people to just ask, in plain language, and get an answer.
Microsoft Research had just opened its doors, loaded with optimism about AI, and that optimism ran straight through to product planning. In tandem with the Decision Theory Group, the work became the Answer Wizard, built by Jack Breese, David Heckerman, Sam Hobson, Eric Horvitz, and others.
An infant form of natural language processing (NLP) for just-in-time help was pretty magical. And made our test users feel powerful.
But that isn't quite the full truth. My interest in AI was driven by something more foundational than searching help files. From the minute I touched a computer, all I could think about was how to embed intelligence in it - something that could automate, create, diagnose, and decide, at a scale no single person could. And I wanted it to think like me or at least allow me to curate whose patterns of thought were worth encoding.
The Answer Wizard was fun, but it was a toy. My very first time meeting Steven Sinofsky was at the end of an interview loop for the Office group he would soon be leading. He bluntly stated "the Answer Wizard is just a weighted keyword search with synonyms." To this day I don't know if he was being factual or testing the thickness of my skin, or both, but I easily conceded the point.
Two lines of research at Microsoft fascinated me much more. The first was MindNet, a project by Lucy Vanderwende and Bill Dolan, that turned Microsoft's NLP parser loose on the world to acquire semantic understanding. The second was Bayesian Networks - roughly, sophisticated expert systems - led by David, Eric, and Jack. One piece to acquire knowledge, the second to reason over it.
There was a belief at the time that social UI - a personified agent - would make computer features feel more approachable. Working on natural interfaces pulled me into its orbit. And agents, generally, were a field in full swing - the idea of a swarm working on your behalf isn't new. But something troubled me about the demos I saw.
I didn't want to command an army of faceless AI agents. I wanted to be an army of one.
Not 100 soldiers.
100 of me.
Skepticism noted, I signed up anyway - it was the boldest, most visible place to take chances with AI. In 1995 we were determined to make Clippy earn his keep, so we pursued the intelligence that would power him - not the version you know. A Bayesian network (the AI hotness of that period) drove a rich and nuanced expert system for Excel that watched you work, inferred your intent, and knew when to offer help. It was more powerful than the Answer Wizard because users could discover new features exactly when they needed them.
I can still see one moment in testing: a user selecting rows and columns, applying formatting, and the bars concluding they might want to turn it into a beautiful chart. Nobody had asked for anything - and the user was delighted to learn about charting.
It was cut near the end. Building, testing, and tuning probabilistic models across every app in Office was expensive, and the UX challenges were a research field of their own. What shipped was the neutered version: the cartoon, minus the brain. The brain survives in Horvitz's demo - which still captures some things LLMs lack, a true measure of confidence and interpretability.
I like to say Clippy was actually a success. The first version of Office that promised to fire him became the fastest-selling Office to date. (I agree, that's a tortured definition of success.)
The lessons: Don't underestimate the engineering and good UX it takes to get probabilistic systems to behave. And maybe don't put cartoons in serious products.
Around that same time, Sam Hobson and I built a Word prototype almost nobody saw: it watched what you were typing, pattern-matched against Encarta (it was like Wikipedia on a CD), searched with Find Fast (our document search technology), and pulled relevant references straight into the document. We were already past help systems - who needs help when the computer can just do the thing?
To my mind, there was always a more interesting question underneath all of it: if Excel could calculate numbers, could it someday calculate ideas?
I asked my first manager at Microsoft this while prototyping the Answer Wizard, the inimitable Joel Spolsky. His answer: "Don't prophesy. Try." So I did. I learned Excel's recalc engine - the machinery that updates everything a change touches - deepened the natural language understanding I'd acquired from personal interest, and joined the Decision Theory Group in Microsoft Research to learn expert systems. I was gathering the pieces. More importantly, I was learning how to design UX around probabilistic systems.
design notes 01–10 It turned out the cartoon was the least of our worries.Useful AI is a deep interface problem, with several user challenges.
- 01 Every answer sounds generic.
- 02 It sounds confident when it's wrong.
- 03 I don't know what to ask.
- 04 The first answer never seems right.
- 05 Its help gets in my way.
- 06 Can I keep some things private?
- 07 The longer we talk, the dumber it gets.
- 08 I rebuild the same prompt every time.
- 09 Checking its work takes longer than doing it.
- 10 I can't trust it to run alone.
In 2000 I founded Conversagent, applying what I'd learned - companies could use it to let customers ask questions in plain language and get answers grounded in their own product data, years before that was a normal thing to want. It combined natural language and expert systems to more deeply understand user intent.
The disappointment was familiar: authoring those language and expert systems was every bit as painful as it had been with the Answer Wizard and Clippy. We spent years making it easier. It never got easy.

Conversagent was many years of continued refinement - over time it became one project of many, and incubated several smaller product ideas. Microsoft acquired the technology in 2017 and I led product incubation in Surface and Windows, under Steven Bathiche, whose Applied Sciences Group was chasing the same questions around natural interfaces and ambient intelligence from the hardware side.
Then the missing piece fell into place. The earlier systems could infer intent and handle narrow, carefully modeled problems - but they couldn't make language itself generally executable. Generative AI made it practical to treat natural language as a formula: something you run.
It was time to build again. My latest app is called Pachinko - you drop in information and everything falls into place, like the ball and the slots. It starts with something deceptively plain: notes and tasks. No sane person builds another notes app - unless the notes are refashioned as something else entirely: personal, high-fidelity context that the rest of the system can act on - and that people can easily audit.
Say you want to live longer and better. You set up a meal plan and an exercise routine, then create a feed of the latest longevity research - all within Pachinko. When something new comes in that actually matters - a study that upends conventional advice, a supplement getting reclassified - your plan gets updated: what changed, why it matters, and what to do about it. It's just there, the way a spreadsheet cell updates itself when you change an input. Same pattern for a stock portfolio, travel plans, competitive research.
As I write this letter, two new facts arrived through Pachinko this morning. First, a widely believed finding about a certain vitamin failed to hold up in a larger study. Pachinko recommended a change to my supplement plan. Second, news of Kimi K3’s performance changed how I was thinking about both product strategy and my portfolio. Viable open-weight models at low cost affect me in several ways; Pachinko connected those implications to decisions I was already making. It did a lot of reading I didn’t have to do, then calculated what each development meant to me. Human judgment still makes the final call. AI speeds me up.
Chat waits for a question.
Pachinko responds to change.
That's across the product: your projects, notes, plans, and saved instructions are the context. New information comes in. The right instructions run against it. The result lands back in your work as a brief, a draft, a warning, a next step, or a revision - a short list of notes to choose from. That is idea recalc. It's also the closest I've come to the army I wanted - a hundred of me, reading the feeds, running my instructions, revising the plan while I'm doing something else.
The obvious objection: proactive help is what got Clippy fired. The difference is where the intent comes from. Clippy watched you work and guessed - and the version that shipped guessed without the brain. Pachinko runs on intent you wrote yourself. And the results land in your work, tucked into collapsible sections - the only announcement is a small counter on the project. Read them, audit them, or delete them in a keystroke.
The original Excel team had a gutsy motto: Recalc or Die. It came from Doug Klunder's intelligent recalc: when one thing changes, do not recompute the whole world. Recompute only what the change actually touches. But above all else, recalculate.
When the facts change, your work changes with them.
That's the bet: goals written in natural language, computed in natural language, so your work responds to change the way an Excel model responds to new inputs.
Is this the final vision? Obviously not. Is it finally the shape of what I hoped to build? Very definitely, yes.
From Excel to Clippy to Pachinko.
It looks like you'd like to recalc or die?
AW-95 / CLP-97 / PCK-26