← Back to Blog
Perspective· 14 min read

The End of the Keyboard Is Real. What Christian Klein Can't See Yet.

SAP's Christian Klein says typing ends within three years. He is right about the destination and wrong about the constraint. The bottleneck is not voice recognition, and the device that replaces your keyboard faces a social problem, not an engineering one.

TL;DR: Twelve days ago, SAP CEO Christian Klein told Fortune that data entry by typing will end at SAP within two to three years, because large language model voice recognition is “super strong” and the only work left is translating speech into business data. He is right about the destination and he has located the bottleneck in the wrong place. Transcription is close to solved. What is not solved is the last mile: deciding when a system has understood well enough to send an email on your behalf, and recovering when it hasn’t. That work is unglamorous, it is not getting faster every model release, and it is the entire company. The second thing Klein’s account leaves out is the device. His timeline assumes the laptop microphone, which is where dictation already lives. The reason to believe the keyboard dies is not that voice is faster than typing. It is that voice is the only input that leaves the desk. That is an argument for a wearable, and the wearable category has a graveyard that has nothing to do with engineering.


The end of the keyboard: voice recognition is solved, the hard part is knowing when to act, and the device that replaces your keyboard will be decided by social permission rather than specs.

On September 16, 2026, Kamal Ahmed published an interview with SAP CEO Christian Klein in Fortune under the headline “The end of the keyboard is near.” A version of the piece ran in January, and the conversation happened at the World Economic Forum in Davos, but the dateline that matters for what follows is twelve days ago.

Inside it is a prediction that reads like a press release for the entire voice industry. Voice recognition from the large language models, Klein said, is “super strong.” Within two to three years, nobody at SAP will type data into a system. They will ask analytical questions out loud, trigger operational workflows, and make entries, including performance feedback and pipeline entries. “The technological capabilities are there,” he said. “It really is now about the execution.”

It is a tidy argument, and the first half of it is right in a way that is easy to miss.

In March 2025, at a press conference in Seoul, Klein had already put a date on it: SAP users would stop manually entering data within two years, and Joule customers would see 30 to 40 percent productivity gains. Eighteen months later he is still on the record, and still pointing the same direction, this time in Fortune. The direction of travel is not in doubt. When the CEO of the company that runs a large share of European enterprise accounting says the typing stops, the typing stops.

We build voice-to-action software for Mac, so we agree with the destination. Our disagreement is about where the work is, and it turns out the answer determines almost everything else, including what the device looks like.

The bottleneck is not recognition, and it will not be recognition

Klein’s framing is that speech-to-text got good and the remainder is translation. Translate voice into business language, translate business language into structured writes, done.

That framing is wrong in a specific and consequential way. It is wrong for the same reason it is wrong to say a self-driving car is nearly done because the perception model is nearly done. Perception is not the constraint. The constraint is what the system is permitted to do with a perception result that might be wrong.

Here is the arithmetic that nobody puts in a keynote. If voice recognition is 95 percent accurate at the word level, and you dictate a 200 word Slack message, you have roughly a one in seven chance of at least one error. That is fine for a draft you will read before sending. It is catastrophic for a system that treats the transcript as a command.

The gap between those two cases is not a better model. It is a permission system. It looks like this:

  • Reads never confirm, because a read changes nothing.
  • Low-stakes reversible actions skip the gate entirely.
  • Anything with real-world consequences confirms, on a card showing the exact action and its exact arguments.
  • Anything the system does not recognize defaults to asking.

That ladder took us a year of daily use to get right, and it has not improved once since, because no model release changes it. Compare that to transcription, which has improved on every shipping model for three years running. If you are trying to predict when voice replaces typing inside a system full of consequential writes, you are waiting on the permission system, and the permission system is not on a model roadmap. It is on an engineering roadmap, and it is nobody’s quarterly headline.

So the two to three year timeline is not obviously wrong. It is just not explained by the thing Klein credits for it. The moment he says “it really is now about the execution,” he has described the hardest part and then declined to look at it, because the execution is not a strategy slide. It is confirmation design, and the field guide to how voice agents actually gate write actions is a few thousand words of unglamorous detail that no amount of ASR accuracy removes.

The device is the argument, not the microphone

Here is the part of Klein’s prediction that survives contact with his own example.

He describes voice at a laptop. The laptop microphone. That is where dictation already is, and dictation has been available, good, and ignored for thirty years. If voice input lives in the machine on your desk, it competes with the keyboard on the desk, and the keyboard wins on anything involving a number, a name, or a symbol. Voice does not beat typing at typing. It beats typing at the thing typing is bad at, which is leaving.

The keyboard is not slow. The keyboard is located. It is bolted to the desk, and reaching it means sitting down. That is the actual constraint voice removes, and it is a constraint no amount of microphone quality in that same laptop will address. You can install a better model in the machine on your desk and you will still be sitting at the desk.

Which means: if you take Klein’s prediction seriously and take it seriously as a mobility claim rather than a transcription claim, it points somewhere specific. An input surface you can carry, that needs no desk, that you can raise to speak into, that works whether you are at the laptop, on the phone, or somewhere with neither. A pendant. A clip. A ring.

The instinct that the next input device is small, worn, and button-gated is not a hardware guess. It follows from the argument. If the win is leaving the desk, then the winning device is the one that is not at the desk.

The wearable graveyard has one failure mode, and it is not engineering

The always-on wearable category has produced a remarkable number of corpses in a short time, and reading them together makes the pattern obvious.

Humane AI Pin. Shipped April 2024 at 699 dollars plus a mandatory 24 dollar monthly T-Mobile plan. The charging case had fire hazard issues and was recalled. Returns outpaced sales, and by October 2024 the price was cut from 699 to 499. Sales stopped February 18 2025, the servers went dark ten days later, and HP bought the software, more than 300 patents, and most of the staff for 116 million dollars. HP pointedly did not buy the hardware. The 699 dollar device customers had actually paid for became a brick.

Friend. A 99 dollar always-on necklace built to listen and keep you company. The launch posters in New York were defaced. Someone spray-painted over the advertising for a product whose core feature was hearing the people standing nearby.

Limitless Pendant. The most direct attempt at the vision: a clip-on microphone, 199 dollars, roughly 20 hours of recording, transcription and summarization per month. It worked, it was popular, it raised more than 33 million dollars from Andreessen Horowitz and Sam Altman, and Meta acquired it in December 2025. Meta then stopped selling it to new customers, discontinued it in the UK, EU, Israel, South Korea, Turkey and Brazil, and is reported to be testing a successor in 2027 rather than shipping one.

Bee. 49.99 dollars, always-on, physical mute button, acquired by Amazon in 2025. Cheapest device in the category and the most commercially successful, which is worth pausing on.

Now read the failure modes. Not one of these devices had a transcription problem. Plaud’s NotePin S, at 179 dollars, records 20 hours on 64 gigabytes of local storage and transcribes reliably enough that it is the recommendation of every reviewer who bothers to check. The hardware is fine. The microphones are fine. The models are fine.

The Friend posters are the tell. The product failed at the point of being in a room with other people. And the pattern repeats in the legal column rather than the engineering column: Limitless required wearers to notify and obtain consent from anyone recorded, a requirement that a device nobody can see makes almost impossible to honor, and Limitless was the product that worked.

David Harris, a former Meta AI researcher and UC Berkeley lecturer, told the BBC that this class of technology is “fundamentally an invasion of privacy and it’s really going to face more and more backlash.” He is describing the constraint, and it is not solvable by a better chip.

So the ring has to be a button, and this is the part we got right by accident

Here is where the analysis lands on our own product, and it is not flattering to say we were prescient, because we arrived at the same place from a different direction and a wearable would stress-test the reasoning harder than a Mac does.

We built mrmr around deliberate capture. The microphone is open exactly as long as your intention is: hold a key, speak, release. Nothing listens between commands, so there is nothing to perform for. The macOS menu bar indicator agrees with you, and we have argued that this agreement is part of the interface rather than a compliance artifact. Our reasoning was about flow. The switching tax research says a knowledge worker needs 23 minutes to fully return to a task after an interruption, and a device that might be listening is a low-grade interruption running constantly in the background.

Put the same design on a ring and the reasoning gets stronger, not weaker, and that is the part worth being clear about.

On a Mac, the microphone indicator is visible and it is on your screen, and you can always see whether it is live. On a ring, you cannot see your own microphone. There is no indicator because there is no screen, no menu bar, no red dot in the corner of a thing you are already looking at. The invisibility that made the ring attractive is exactly the property that makes always-on capture unworkable on a ring.

So the ring does not argue for ambient listening. It argues harder for the button, because the button becomes the only trustworthy signal available. A press you can feel is the privacy indicator. You can hold a ring and know with certainty it is recording, and you cannot say that about a device you cannot see. Hold-to-talk stops being a UX preference and becomes the interface, because on a form factor with no display it is the only honest way to answer the question every person in the room is silently asking.

That is a real design constraint, and it points the same direction as the consent problem above. The device that replaces the keyboard has to make its listening state unmistakable, because the thing standing between it and adoption is not accuracy, it is being the person in the room who cannot explain what the thing on their hand is doing.

We will get this wrong in ways we cannot predict from a Mac app, and the wearable form factor will find them.

What mrmr is actually building is the layer, not the thing

Being precise about this, because the temptation in a piece like this is to imply hardware that does not exist.

mrmr today is a Mac application. It is the voice-to-action layer for your machine: dictation that lands at your cursor in the app you are already using, actions that execute across Slack, Calendar, Reminders, your scripts, and your files, and a sub-agent system that takes bounded work and reports back. It is not a ring, it is not a wearable, and nothing in this post should be read as announcing one.

The architectural position is the one that matters for where the device goes. The agent is a layer. The input surface is interchangeable. Today the input surface is a key you hold on a Mac. The shape of the thing you are building makes it possible for the input surface to become a worn device, a phone, a headset, or something nobody has thought of yet, without the agent changing. What does not change is the part that is hard: the permission ladder, the context resolution, the local execution across apps that only you have access to.

That last one is worth dwelling on, because it is the part the giants are least likely to build. An agent that can read your email, your Slack, your files, and your screen and take actions with them on your behalf is a security liability unless it is scoped to a single person’s machine. Cloud agents cannot do that. They run somewhere that is not your laptop, and the value of the thing is precisely that it has access to the things your laptop has access to. The reason the layer is worth building locally is not privacy theater, and we have written about exactly what a voice agent sends to the cloud and what should stay local. It is that local execution is what makes consequential actions possible at all.

Which also tells you where the wearable pressure comes from. The moment a small device is asking your Mac to do things, it is asking across a trust boundary, and the design question stops being “is the audio good” and becomes “what is this allowed to do, and how do I know what just happened.” That is the same question as the confirmation ladder. It is the same problem at a different distance from your face.

Why this is defensible against Meta and Apple

The giants are coming, and it is worth being clear about what we are and are not competing with.

Meta sold more than 7 million Ray-Ban smart glasses in 2025 and holds roughly 82 percent of the smart glasses market. Per a leaked internal memo from Alex Himel, the company’s VP of wearables, Meta plans to test an AI pendant in 2027, expand the glasses line, and launch an enterprise subscription called Wearables for Work. Apple, per Bloomberg reporting in February 2026, is accelerating three wearable products including an iPhone-dependent pendant.

The counterweight is in their financials. Reality Labs lost 4.03 billion dollars in the first quarter of 2026 on revenue of 402 million.

So the wearable future is real, well funded, and probably will happen. The question is not whether Meta can build a ring. They have shipped glasses at volume, which is harder in some ways than a pendant, and their privacy problem is a public-relations problem they can spend money on.

The defensible ground is different, and it is the ground the corpses are sitting on. A device that only records is a commodity, and it already is: 179 dollars for a reliable transcriber, 49.99 dollars for an always-on one. A device that acts is a different product, and acting is where the whole difficulty lives, and difficulty is where a small team has an advantage over a large one that is trying to protect a hundred million people in smart glasses.

Nobody at a company shipping 7 million units a year wants to own the question of when their hardware should send an email on a user’s behalf. That is our question. We have been answering it for a year, in public, and our answers have a documented cost: the 23 minute attention reset, the half-second discard on an accidental press, the reads that never confirm. The hard problem around a camera pointed at your colleagues is a different problem from the hard problem around an agent holding your Slack credentials, and they are better at the first one.

What would change our mind

The strongest version of this argument is not that voice replaces the keyboard everywhere. It is that voice replaces the keyboard for a specific, well-defined class of work, and that class is larger than most people assume but smaller than Klein implies.

Voice wins where the work is short, mobile, conversational in origin, and consequential enough that a modal dialog would cost more than the error. Sending a message that responds to something you just heard. Filing a note while walking. Triggering a script. Dictating into a field. That class is real and it is growing every model release.

Voice loses where the work is long, dense, numeric, revision-heavy, or where you need to see the state of the system to know what to say next. Nobody dictates a spreadsheet. The keyboard survives in code review, in financial modeling, in anything where the feedback loop is visual.

And the prediction that concerns us is not that voice wins. It is the discovery cycle. The first 312 years of the keyboard were the typewriter, and the typewriter did not replace writing, it replaced a professional. When the interface finally does displace the incumbent, the displacement arrives all at once, in a product cycle, not gradually. Klein’s own framing is a warning about this: “business cannot just change the software.” The software that replaces the keyboard will be the one that already had the layer built, sitting on your machine, when the device arrives.

That is what we are building. The layer that a small worn device, or whatever comes next, will need in order to touch your Mac, your phone, and your files without becoming the most dangerous object in the room.

The end of the keyboard is real. The harder part is that nobody has yet built the thing that is allowed to replace it.


Christian Klein told Kamal Ahmed for Fortune on September 16, 2026 that voice recognition from large language models is “super strong” and that SAP will stop data entry by typing within two to three years. We think he’s right about the destination and wrong about the constraint. Transcription is not the bottleneck. Authorization is.

Private beta

Get private beta access

Book a short setup call or join the invite list for Agent Mode access.