Newest first, condensed for someone deciding whether to install or update. The full notes and the download for each version are on GitHub.
2.1.2
Current
Two forwards at once: each Agent run now stays scoped to its own working directory instead of wandering, and the dictation history finally updates when you correct something in place — the correction window writes the corrected full text back into history and memory, instead of forever learning the exact wrong version you just fixed.
New
Each Agent run now starts from its own working directory. Previously an agent's default working directory was your whole home directory, so it could search anywhere. Now every agent task can carry an explicit working directory: `bash`/`python` default their `cwd` to it, `glob`/`list_dir`/`grep` default their search root to it, and relative paths resolve against it — while absolute paths and `~`-prefixed ones stay untouched.
History and memory now follow after an in-place correction. When you fix a misheard word through the correction window in Direct/Tidy mode, the app only pastes the replacement back into the target app — but history and the recent-activity context remember the whole pre-correction sentence, which is exactly the text you just corrected. Now the corrected whole is rebuilt, purely locally and by rules, and written back: only when the corrected span occurs exactly once in the delivered text; an absent or repeated selection rewrites nothing, never a guess.
Fixed
`grep` with no path no longer wanders off into ~/Desktop and ~/Downloads. It now falls back to the current run's working directory like the other first-party tools.
Known gaps
An in-place correction is a purely local rule-based rewrite, and it is only safe when the corrected span occurs **exactly once** in the whole delivered text. If the same typo appears twice or more, this release leaves history alone — with only text and no live selection range from the target app, it cannot tell which occurrence you meant, and prefers not to guess.
The working directory is per-run, but there is no dedicated place in the app UI to set it for the Agent yet. It is currently passed in by the caller through the API; the default remains your home directory.
2.1.0 moved Skills and Agents into the interface. 2.1.1 fixes something nobody had caught: every time "Clear local data" actually deleted something, it told you it had failed.
Fixed
"Clear local data" reported a failure on every successful reset. The speech service answers a deletion with how many rows it removed — a number — while the app decoded that field as a boolean, which always throws. So a reset that really cleared your history was recorded as an error; the only time it appeared to succeed was when there was nothing to delete. Both shapes decode now, and a genuine network or HTTP error is still reported as one.
Agent run logs had no cap and no way to clear them. Every agent run left a file in run-logs/, growing without bound since day one — and nothing in the product ever read it; a run in progress reads its steps from memory. The newest 50 are kept now, and "Clear local data" clears them too: they record what you asked the agent to do, which is input history.
The same shape, a second time: clearing run logs failed as a whole if any single file couldn't be deleted — say a concurrent run's own pruning removed it first — turning "deleted most of them" into a reported failure. Each file is handled on its own now, and the count reflects what actually went.
Known gaps
The statistics panel's figures still come from that append-only audit log, which "Clear local data" still doesn't touch. So the numbers stay put after a clear — 2.1.0 said as much, and 2.1.1 doesn't change it.
This release moves two things out of the filesystem and out of a one-shot panel and into the interface: Skills and Agents are now visible and editable, and every step an Agent took stays with the conversation. The cost is stated up front — upgrading clears your existing conversations.
New
Settings → Engines → "Skill 与 Agent" is a new management page. Extending the Agent used to mean dropping files into ~/.opentype and having no way to see which ones actually took effect. Now the three sources — built in, your own, and whatever is read from your Claude Code directory — appear in one table you can create and edit in place. When a name collides the built-in wins, and your copy is labelled as shadowed rather than silently ignored, which is what used to happen. The files themselves are unchanged and still readable by Claude Code.
Agent and Ask now store their execution steps alongside the conversation. That log used to live only in the panel of the run that produced it — close it and it was gone, and opening an old conversation showed nothing. Now any past conversation expands to show which tools ran and in what order. The cost is that upgrading clears your existing conversations — the storage shape changed and we chose to rebuild rather than migrate. Dictation history, the entity dictionary, and what the app remembers about you are untouched.
Fixed
The "Agent tools" page listed 10 built-in tools. In reality the Agent had already gained six more — write a file, edit one, move one, put one in the Trash, find files by name, and read back dictation history — and the page never caught up. All 16 are listed now, with the 8 that leave something changed on your machine marked as having side effects. The question people open that page to answer is what the Agent can actually do, and the answer it gave was wrong.
The explanatory copy is gone from the interface — where files live, what order things are discovered in, how the index works. None of that earns space in Settings. What stays is state and values, keyboard hints, and the consequence line a destructive action shows in its own dialog.
Settings sub-pages drew their title twice — once from the shared page chrome, once by the page itself. "Agent tools" had shipped that way from the start, and the new "Skill 与 Agent" page inherited the same mistake. Both show one title now.
Known gaps
An agent definition still has to be named at the start of the task ("use writer to draft an email"). The main agent won't pick one for a subtask on its own.
A step log in a past conversation has no elapsed time — the steps were stored, how long the run took was not.
The four numbers in the statistics panel are computed from an append-only audit log, and nothing in the product clears it: "Clear local data" clears dictation history, not that. So after clearing your data, the statistics stay where they were.
A Skill read from ~/.claude whose description uses a YAML folded block (a bare > after description:, text on the following lines) comes through as a lone ">". That description is the only thing the model sees when deciding whether to use the skill, so such a skill effectively never gets picked. Skills you create inside OpenType are unaffected.
The headline this time isn't a fix — it's a deliberate change to a stated promise: Ask and Agent now carry recent conversation as context, dictation included, with no switch to turn it off.
New
Ask and Agent requests now automatically carry the last 10 turns as context, and all three modes count toward that window — dictation included. A question or task can build on something you just dictated, without repeating it. There's no switch to turn this off.
This isn't the 1.0.0 bug come back: that one was dictation quietly reaching the cloud through memory consolidation, against the promise that dictation never touches a model. This is a deliberate, stated decision. Audio is still recognized on-device and deleted right after — what changed is that the recognized text gets reused as context, not that dictation itself now needs the network.
Fixes for things the product claimed to do and didn't — a follow-up card wired to nothing, a recording with no visible way to stop, and four more like it.
Fixed
The Ask result card's follow-up controls — the text box, the microphone button, "say more," "rephrase" — were wired to nothing. Typing an answer and pressing Return did nothing at all.
A recording started by that microphone button had no visible way to stop: the button was hidden under a pill reading "release to finish," a key you never pressed. It now says "click to finish" instead.
Saved MCP server changes now take effect immediately, and only the server you actually edited reconnects — changing one no longer drops every other server's tools for a few seconds.
"Reset input history" now actually clears the debug log that kept the first ~200 characters of every question and task. It used to survive a reset the dialog called irreversible; that log also got real file permissions and a 5 MB cap.
A bad memory-consolidation run can now be rolled back, most recent first, one run at a time.
The "average wait" stat no longer counts time spent reading and editing in the Review panel as system latency, and dictation and Ask times are no longer averaged into one number.
Known gaps
Anthropic as a provider, remote Whisper, a real MCP server, and voice correction with an actual microphone are all still unverified end to end — each needs a real key or real hardware to exercise.
A save-time MCP connection failure still doesn't show up until you close and reopen the panel.
A product that had stopped telling you what it was doing. This release makes it say so, and shows what it has quietly learned.
Fixed
The silent first-run download of the ~460 MB speech model — which looked like a hung app — now shows progress. You can talk while it downloads; your words wait for the model instead of being dropped.
Dictionary learning is now visible: a rewritten transcript says how many words it changed and which, a correction says what it taught, and a wrong guess can be undone.
The 27-language transcription picker used to change nothing about the result. It now actually reaches the recognizer.
An MCP tool server that silently timed out at startup used to look identical to one that worked. The row now says which happened, and why.
New
Model choice moved into Settings, instead of only being set through an environment variable a packaged app can't set.
Update notifications, launch at login, an English interface, and history you can search, delete one entry at a time, and export. Terminals stopped receiving auto-inserted dictation, since a pasted newline there is a command.
Dictation that learns your terms, a main window rebuilt around what you came to do, and agent tools you can actually see and configure.
New
Your entity dictionary now biases speech recognition before it transcribes and rewrites the result afterward, and a correction teaches the dictionary instead of being thrown away.
An optional "tidy" cleanup: filler words and spacing fixed by plain rules, no model call, so it stays as private and as fast as plain dictation.
The main window was rebuilt: a sidebar and one main view instead of five equal tabs, with conversations you can continue instead of restarting.
The agent's built-in toolset is listed in Settings, MCP servers can be added and removed from the UI, and commands that delete or overwrite files now ask before running.
Fixed
Mode switching didn't work at all on some hotkey presets, one bad MCP server could stop the whole app from starting, and dictation was quietly reaching the cloud through memory consolidation despite the promise that dictation never touches a model.
Worth being honest about: this was a large rewrite, and not all of it was hand-run. The rebuilt window and mode switching were run and seen working on a real Mac. Voice correction, the Agent tools screen, and connecting a real MCP server were covered by automated tests but not exercised end to end by a person — those are the areas to be suspicious of.