更新日志
从 GitHub Releases 拉取的版本说明。
v1.0.282
2026年10月9日 · v1.0.282在 GitHub 查看 ↗✨ New: Claude Haiku 5.5
Claude Haiku 5.5 (
claude-haiku-5-5) is in the Claude model picker above Haiku 4.5. It supports effort fromlowtomax, plus ultracode and ultrathink, and defaults tomedium. It costs $0.10 / $0.50 per MTok with cache reads at $0.01. Prompts over 100K tokens bill at 5x, and Token Stats only tracks totals per model, so the cost it shows for Haiku 5.5 is a lower bound.Engine pins: Claude Agent SDK 0.3.285 → 0.3.293 (Claude Code 2.1.293), Codex 0.159.2 → 0.161.0.
v1.0.281
2026年10月1日 · v1.0.281在 GitHub 查看 ↗🐛 Fix: Comments button only lights up when there is something to read
The comments button next to the chat input used to be tinted all the time. It now uses the same muted color as the notes button until the project has at least one code comment, and picks up the tint as soon as one is added. Deleting the last comment mutes it again.
📚 Docs: Running the agent on any model
The engines page now opens with a direct answer to "how do I use Claude Code with another LLM". A new section compares overriding Claude Code's environment or putting a gateway in front of it with picking an engine per tab, and covers what the Built-in Agent gives up.
🌐 Site: Two new posts and a sharper headline
Two new blog posts: Claude Code with GLM, Kimi and DeepSeek — No Env Vars, and a side-by-side comparison of a terminal CLI workflow with the OpenCockpit workbench, including when to pick which. Page titles, the homepage lead, the social card and the footer now say "Open Claude Code GUI for any LLM".
📦 Misc: npm description and README
The npm description, keywords and README headings now use the same any-LLM wording as the website.
v1.0.280
2026年9月30日 · v1.0.280在 GitHub 查看 ↗✨ New: Every running terminal, from the sidebar
Terminals now outlive Cockpit restarts, which made a dev server left running in another project easy to lose track of. A Terminals row under Scheduled Tasks in the sidebar shows how many commands are live and opens a board grouped by project. Clicking a row switches to that project, swipes to Console and selects the bubble; each row can also stop its process. A stopped command stays in its bubble as finished, the same as stopping it from the bubble itself.
✨ New: Claude Sonnet 5.5 and GPT-6.1 Sol
Claude Sonnet 5.5 (
claude-sonnet-5-5) is in the Claude model picker, with effort up toxhighandmax, and ultracode support. It defaults tomediumeffort and costs $2 / $10 per MTok. Opus 5.5 stays the Claude default.GPT-6.1 Sol (
gpt-6.1-sol) is codex's new default and is now Cockpit's default Codex model, atlowreasoning.✨ New:
/gokeeps a decision log/gonow works from the agreed spec. It verifies each slice by running it, retries a failing slice at most three times before marking it blocked, and stops only for choices that are costly to undo. Every decision the spec did not cover is logged as one line to~/.cockpit/skills/go/notes/<project>-<feature>.md, linked from the recap. The notes are cleaned up after 30 days without changes./apis gone, since its decision log now lives in/go./qa,/fxand/exnow produce fixed, skimmable documents./qawrites a lean PRD,/fxa root-cause analysis with a cited evidence chain, and/exan answer-first analysis. Follow-up rounds show only what changed. The/crfindings index now renders as a list instead of one run-on paragraph.✨ New: Baseline shifts hidden in the per-call diff list
A turn that starts with a branch switch used to open the per-call diff viewer on a 200-file call made of other people's commits. Those leading baseline-scale calls are now hidden, as the aggregate view already did. A banner says how many were hidden and shows them again in place.
🐛 Fix: Terminal reruns and first-frame size
Rerunning a dev server that took longer than 200ms to exit could leave the new run marked finished with no output, because the old run's late exit was routed to it. Each run now has its own session. A new PTY also used to miss the bubble's first resize and start at 120×30, so the zsh right prompt wrapped and stray
%lines appeared. It now opens at the bubble's actual size.This update leaves the pty-host untouched, so terminals that are already running survive it.
🐛 Fix: Scheduled tasks under concurrent writes
A read that landed during a scheduled-task status write saw an empty file and failed with "not valid JSON, refusing to overwrite it". This broke the task panel, and at boot it left no timers armed until the next restart. Reads now take the same file lock as writes.
📚 Docs: Quickstart ends with
/crThe feature-work quickstart now runs
/crbefore the manual review in Explorer, and the built-in skill count is down to 12.Engine pins: Claude Agent SDK 0.3.283 → 0.3.285, Codex 0.158.0 → 0.159.2.
v1.0.279
2026年9月28日 · v1.0.279在 GitHub 查看 ↗✨ New: Terminals survive a Cockpit restart
Terminal commands used to be children of the server, so every update or restart killed them — a dev server started in a Console bubble died with each
npm i -g @surething/cockpit. They now run in a separate pty-host process (cockpit-ptyin your process list), which the server talks to over a local socket. On boot the server adopts the sessions that are still running, together with the output they printed while it was down; commands that ended in the meantime come back with their real exit code. Both PTY and pipe mode are hosted.cockpit stop # shuts the pty-host down too — this ends your terminalsUpdate and restart leave the host alone. The one exception is a release that changes the host's own code: the new server then replaces the old host, which ends the terminals it was running. The host exits by itself once it has no sessions and no clients.
Terminal output is also persisted more reliably. Finished PTY bubbles used to come back blank after a refresh, and they now keep their output. Deleting a bubble or clearing a tab removes its output file, and orphaned output files are swept at startup.
\r-only progress bars in pipe mode no longer grow without bound.An Ollama server started from the UI survives restarts too. It used to be killed along with the server's other child processes on every stop, update or restart.
✨ New: An in-page folder picker
Open Folder used to show an OS dialog launched from the background server process. Each platform broke it differently: macOS 26 would not bring it to the foreground, so paste never reached it; on Windows it opened behind the browser; and on Linux under ssh or systemd there was no display to show it on. Opened from another device, it appeared on the server's screen. The picker now lives in the project browser. It has an address bar with completion: type or paste a path, Tab/Enter goes into a directory, Cmd/Ctrl+Enter opens it, Backspace goes up and Esc returns to the project list. Git repos and hidden directories are marked.
🐛 Fix: Wide tables no longer flip the view
When a horizontal scroll reached the edge of a wide table, the rest of the swipe, momentum included, passed through and switched panels. A swipe now belongs to whichever pane it started over, until the wheel goes quiet. The trade-off is that a swipe over a code or diff pane with long unwrapped lines no longer switches views or dismisses the diff column. Use the top bar or the diff's ✕ instead.
🐛 Fix: Independent-task turns no longer hide the history
While an independent (no-history) turn was running, opening that session in another tab could replace the whole transcript with the single in-flight turn. The view then jumped to the top once the turn finished. Readers now see the stashed history and the running turn together.
🐛 Fix: Long quick instructions stay on one line
In the quick-instructions editor, a long instruction used to wrap and look like several separate records. Each instruction now stays on one line and scrolls horizontally, and the editor can be resized again.
📚 Docs:
/crrestructuredThe built-in
/crprompt now spells out how the review runs. The main session dispatches independent reviewers and assembles the report, but never reviews anything itself. By default there are two reviewers, one static and one dynamic; tiny diffs get a single reviewer, andhardfans out one reviewer per slice. Each finding carries a severity and an origin (introduced, activated or pre-existing), and the report has a fixed format.Engine pins: Claude Agent SDK 0.3.280 → 0.3.283, Codex 0.156.1 → 0.158.0.
v1.0.278
2026年9月23日 · v1.0.278在 GitHub 查看 ↗✨ New: Output styles
A global library of named output styles, picked per session from a new toolbar picker. Styles are edited in one textarea — a
# Nameheading per style, the body underneath taken verbatim — and stored in~/.cockpit/output-styles.json. None is the default and injects nothing.The session stores only the style id; the text is resolved on every dispatch. An edit therefore reaches the next turn of every session using that style, and a scheduled task reads its session's selection at fire time rather than when it was scheduled.
Each engine receives it in its own place: appended to Claude Code's system prompt, sent as developer instructions to Codex (reconciled on resume, since Codex ignores that parameter there), and added as a final section for the built-in agent loop. An unchanged style renders byte-identical, so Claude's prompt cache still hits after a switch back.
This also fixes Claude sessions that had been running without Claude Code's system prompt: when no system prompt was given, the Agent SDK sent an empty custom prompt instead of the
claude_codepreset. The preset is now always passed.✨ New: Claude Opus 5.5 and GPT-6 Sol / Luna
Claude Opus 5.5 is the new default Claude model, at
mediumeffort — Claude Code's own default for it, and the one model in the lineup that does not default tohigh. It supports xhigh, ultracode and fast mode. GPT-6 Sol is the new default Codex model atmedium, with GPT-6 Luna alongside it; labels, effort sets and defaults come from the catalog Codex ships.Both pickers are trimmed to the current lineups. Claude drops Opus 4.8, 4.7, 4.6, Sonnet 4.6 and Fable 5, keeping Opus 5 as a fallback; Codex drops GPT-5.6-Terra and GPT-5.6-Luna, keeping GPT-5.6-Sol. A session pinned to a removed model keeps running on it, and its picker now shows the model's name (
Claude Opus 4.8) rather than the raw id.The 200K / 1M context toggle is gone. Every model left in the picker runs a native 1M window, so it was offering a choice that changed nothing.
Engine pins: Claude Agent SDK 0.3.278 → 0.3.280, Codex 0.155.1 → 0.156.1.
✨ New: Markdown preview in history diffs
The history tab's compare mode and commit detail now offer Preview on Markdown files, opening the interactive preview on the "after" content — with comments and send-to-AI. ESC closes only the preview.
🐛 Fix: Git status in worktrees, JSON preview in compare mode
A commit made inside a git worktree only rewrites the branch ref in the main repository's
.git, which the watcher did not look at — so the Changes tab stayed stale until something else refreshed it. It now follows the worktree's common git dir.In compare mode, the Readable button for JSON files did nothing until you switched to the status tab. It now opens the preview directly.
🐛 Fix: Command tags no longer change with the UI language
Resolved command tags read
[主会话·qa]next to[subagent·cr]for a Chinese UI, and scheduled tasks fell back to English regardless. The tag is model-facing metadata, not UI copy, so it is nowmain/subagentin every language.🌐 Site: Long-term memory is just a directory
A new blog post, in English and Chinese, on why Bots keep memory as a plain directory the model explores like a codebase — with
BOT.mdas the index — instead of a vector database. The/trydemo, which failed to start after thecockalias removal, is fixed too, as is an install snippet on the site that still used the old name.v1.0.277
2026年9月21日 · v1.0.277在 GitHub 查看 ↗✨ New: Read a whole turn as one diff
The chat diff viewer only ever answered "what did this tool call do". Reading a turn meant clicking through every call and summing the overlaps by eye. An aggregate toggle now collapses the range into a single diff, built over the shadow snapshot repo from the first call's parent to the newest one.
It is a net diff, and that is the point: a file edited five times appears once, and a file created then deleted inside the range does not appear at all. The count therefore drops — 41 per-call entries can be 12 distinct files — so the header labels the aggregate rather than letting that read as lost data.
A turn that opens with
git checkout -b x origin/mainused to start the range on a call that rewrote 200 files and 17k lines of other people's commits. Nothing in a snapshot says "this was a branch switch", so size stands in for intent: oversized leading calls are left out of the starting point. That is a guess, so it is never silent — the meta bar states how many calls were dropped and toggles them back in.Two interaction changes ride along, both about closing the viewer. The header controls move to the left, macOS-window style, since the pointer already lives over the left half of the diff and a ✕ in the far corner charged a full-width trip for the most common action. And the diff column can now be swiped right to dismiss: the gesture draws itself, the column trailing the swipe at half its travel before flying out or springing back.
✨ New: Branch compare counts work you have not committed yet
The history tab's branch compare ran
git diff <base> HEAD, so a branch whose work was still in the working tree rendered "no changed files" while a pile of fresh edits sat on disk. Between that view and the status tab, nothing answered "what did this branch change in total" — the question an agent-driven session asks most.The HEAD pill in the compare header is now a two-segment scope switch:
- worktree — committed + staged + unstaged + untracked (the new default)
- head — committed work only, the GitHub "Files changed" diff
Untracked files are merged in from
git status, sincegit diffnever reports them. Both modes now anchor the old side at the merge base; the previous two-dot spelling reported commits made on the base branch since branching as if this branch had reverted them. A project opened at a subdirectory of its repo also read working-tree paths against the wrong root and returned an empty diff.✨ New: A Bot session says so, wherever it is listed
A Bot session used to be indistinguishable from one a person opened. The mark now rides on chrome that already exists rather than adding a glyph: the number chip keeps its number and changes shape — a robot head, masked over the same status wash — and the running line gets a badge. Detection reads the
@nameout of the dispatch line once, server-side, and it travels with the payload that already answers which engine.An
@botline sent from a session that is already sitting in that Bot's directory now runs in place instead of spawning a child. Delegating exists to give a Bot a session whose files are its own; a session that already has that bought a cold start and a second transcript and nothing else. Matching is exact after realpath, so a subdirectory of the Bot does not match and a repo root containingbots/cannot swallow every@botline in the project. A delegation with no explicitenginealso inherits the engine of the session that asked for it.🐛 Fix: Ollama runs the model your machine actually has
Ollama ids are whatever the machine pulled, so any id baked into the source is a guess that
ollama pull/rminvalidates — and one delegation could produce three separate 404s from Cockpit rather than from the model. Model resolution now happens against the machine: the request's own model (a baregpt-osscompleted to its one installed tag), else the model this session already ran on, else the model the last run used, else the catalog's first entry. An empty or unreachable catalog is now an error naming the server URL and the fix, raised before the run starts rather than landing in a transcript you then throw away.A new ollama chat opens on the model last used instead of on "Select model", and the picker and the engine read the same resolver, so they cannot answer the question differently.
🐛 Fix: A scheduled task follows its fresh session immediately
When a task's resume target is gone it starts a fresh session, and that session's id used to be written back only after the run finished, and only on the success path. So for the whole turn the task pointed at a session that no longer exists — the board opened an empty transcript while the real one streamed under another id — and a run that timed out left the dead id on disk, which is self-reinforcing: the next round starts yet another from-scratch session, and a from-scratch run is the one most likely to fail again. The task now binds on the id the engine announces seconds in. The per-run deadline also rises to 60 minutes; it is a runaway guard, not a latency budget.
🐛 Fix: Quick instructions can span several lines
The outline editor's record separator is the line break, so a multi-line instruction had no way to spell itself — and the round trip was already lossy, silently tearing such an instruction in two on the next save. Real newlines are now encoded as a literal
\ninside the textarea and decoded on the way out, so storage, the API and the send path keep carrying real newlines.📦 Misc: The
cockalias is gonecockis an English profanity, which makes the short alias awkward in shell history, CI logs, documentation and talks. Cockpit ships a single name now:cockpit # production server cockpit-dev # dev serverIf you had
cockin a script or an alias, switch it tocockpit— same 237-line implementation, now living inbin/cockpit.mjs. The lockfile carried a stale entry pointing at the deleted file, which would have leftnpm ciwith a dangling symlink.Engine pins move with this release too: Claude Agent SDK 0.3.268 → 0.3.278, Codex 0.154.0 → 0.155.1.
v1.0.276
2026年9月18日 · v1.0.276在 GitHub 查看 ↗✨ New: Bots — a folder you address as
@nameA Bot is a plain directory with a
BOT.mdat its root. Register it, then start a line with@name: that message runs as its own session, carrying the Bot's files as long-term context.@reviewer look at the auth changes on this branchCockpit does three things and no more — resolve the name, rewrite the message, expose the delegate and status endpoints. It never reads, writes or locks a Bot directory, so a Bot stays an ordinary folder you can edit by hand, keep in git, or hand to someone else.
The command surface is now split by prefix:
/verbruns in the main session,/@verbis delegated,@nameis a Bot. Create one with/bot. Writes are explicit and locked — a Bot's files change only when you ask it to remember, update, correct or forget, under<bot>/.locks/write, whose owner file records the run so a stuck lock is decided by asking whether that run is still alive rather than by a timer.The full design, including the failure mode each rule exists to prevent, is in
docs/BOTS.mdand at Bots.✨ New: Share a Bot with
@name export to <path>A Bot's working directory is the wrong thing to share. It holds the write lock, review reports, your own memory, and paths that exist on exactly one machine — hand it over and all four go with it, usually into a public repository where the mistake cannot be taken back.
@reviewer export to ~/share/reviewerExport builds a new directory instead:
BOT.mdrewritten,identity/andskills/carried over, everything else recreated as empty scaffolding. It is a whitelist rather than a blacklist, so a Bot that grows anotes/next month does not quietly start shipping it. The report ends with the Skill names the copy references, which is the whole installation requirement for whoever receives it. What happens to that directory afterwards is your business — Cockpit neither publishes nor registers it.A Bot's Skills table now names its tools instead of locating them: a registered skill's name, a Bot-relative path for one the Bot grew itself, an absolute path only when it is neither — flagged as machine-local when it is written. That is what makes a shared table mean anything. The recipient registers the skill under the same name, their own copy in their own location, and the row resolves for them as it did for you. No path could do that, because the only paths two machines share are the ones neither of them chose.
✨ New: Built-in Bots live in
~/.cockpit/botsA built-in Bot used to run from the install root — root-owned under
npm i -g, replaced wholesale on every upgrade, shared by everyCOCKPIT_HOMEon the machine. The panel sent you there anyway, so an edit was either refused outright or silently deleted by the next upgrade.The shipped directory is now a seed. Each built-in is installed into
~/.cockpit/bots/<name>, and that copy is the Bot: the panel card, the folder button and@namedispatch all follow the same path to a directory you can open and edit. Install tracking is per file, so editing the persona keeps your edit and still takes laterBOT.mdfixes — where before, one edit froze the whole Bot at the version you touched it. Deleting the folder resets it.cockpit-helper, which answers questions about Cockpit by reading opencockpit.dev live, now keeps memory as a result — of the person, never of the site. Which install you run, how you want answers, corrections you made, what you asked it to follow up on. Not a remembered route through the docs, which is how it would stop reading the site and start guessing at it.🐛 Fix: A malformed registry is no longer overwritten with emptiness
bot.json,skills.jsonandscheduled-tasks.jsonare files you are invited to edit by hand, so a stray comma is a realistic state rather than a hypothetical one — and it read as "empty", after which the next write persisted that emptiness. A corrupt file now fails the write instead of replacing your data behind a success toast.The scheduler's boot path is the deliberate exception: it catches, logs that no task will fire, and refuses to save while in that state. Propagating the throw there would have traded silent data loss for a Cockpit that will not start.
🐛 Fix: Cockpit's own
PORTno longer leaks into your projectsThe launcher exports its listening port as the generic
PORT, which Next, Vite and most dev servers read as their port. It is now stripped from every process spawned into a project.COCKPIT_PORTstays, for CLI bridges that want it, and a custom port still survives a restart or update.Each engine also exports
COCKPIT_CWDnow — the directory the session was started with. Everything keyed on cwd (delegation, the Bot write-lock owner, status lookups) previously had to work it out, and a session running in a subdirectory worked it out wrong: it saw the repository root's.git, called that "the project", and wrote records nobody could look up.pwdis not the answer either, since it follows anycdthe turn has made.✨ Chat and Git polish
The branch compare header now carries summed +additions / −deletions on its right, next to "N files changed", styled like the per-file counts. The
<engine> running <elapsed>line gets an orange shimmer sweeping across it, matching the running spinner —@supports-guarded, and off underprefers-reduced-motion. Notes now sit before Skills in the sidebar, and the Git change-classification badges line up.🌐 Site: SEO metadata and breadcrumbs
Per-page canonical and OpenGraph metadata for docs and blog, breadcrumb markup, a corrected
robots.txt, and a validation pass in the site build. Docs<lastmod>dates were also re-derived from git — 22 of them had drifted, which matters because a sitemap whose dates are learned to be wrong gets trusted less.v1.0.275
2026年9月14日 · v1.0.275在 GitHub 查看 ↗✨ New:
/ssfinds a past session from one sentenceYou remember what a conversation was about, not which project, engine or week it happened in.
/sstakes that half-memory and finds the session:/ss the session where we worked out CSRF on the local APIThe agent does not search your sentence verbatim. It expands it into the words that would literally appear in that conversation — both languages for technical topics (
跨站/CSRF), synonyms, the terms the assistant would have used — searches every session Cockpit can read across all projects, all engines (Claude, Codex, DeepSeek, Kimi, GLM, Ollama) and all dates, reads the snippets, and replies with 1–3 candidates. Each carries a session link; clicking it switches to that project and opens the session in the Agent panel.Behind it, Cockpit keeps a text-only copy of each session's prompts and replies under
<data-dir>/search-corpus, updated incrementally, and searches it with ripgrep, so two-character Chinese words like快照match. The first search on a machine takes a few seconds while that copy is built.✨ New:
/dlhands work to another session without waitingHalfway through a task you notice work that belongs somewhere else — a flaky test in another repo, a job better suited to Codex.
/dlstarts it there and returns at once:/dl have codex fix the flaky date test in the api projectThe agent writes a self-contained brief (the child sees none of your conversation), starts a new session in the target directory on the chosen engine, and repeats the receipt — engine, directory, link — in its reply before carrying on. The child is an ordinary session: open the link to watch it or take over. Later, ask how it went; the agent finds the receipt, in this conversation or through
/ss, and reportsrunning,done,failedorincompletealong with the child's last reply.Nothing is stored server-side — status is read from the child engine's own transcript. At most 4 delegated sessions run at once (
COCKPIT_DELEGATE_MAX); past that the request is rejected rather than queued.✨ New: A Host / Origin check on every request
A web page open in your browser could previously reach a local Cockpit through DNS rebinding or a cross-site request. Every HTTP request and WebSocket upgrade now passes a check first: a request arriving over loopback must address Cockpit by a local host name, and a state-changing request or WebSocket upgrade that carries an
Originmust be same-origin. Rejections are403 Forbidden.curl, the CLI and skills send noOriginand are unaffected.If you reach Cockpit through a tunnel (ngrok, cloudflared, …), this changes behaviour. The tunnel connects over loopback with its public hostname in
Host, so it is now refused unless you allow that name or run in token mode:COCKPIT_ALLOWED_HOSTS=my-box.ngrok.app cockpit # no token: anyone with the URL gets in cockpit --token my-secret-value # token mode: the tunnel must forward X-Forwarded-ForKeep the tunnel's original
Hostheader; if it is rewritten tolocalhost, POSTs and WebSockets get 403.🐛 Fix: A one-time task that came due during a restart now runs
Scheduled timers die with the process, and a one-time task whose moment passed while Cockpit was down was marked completed without ever running — and then showed as failed in the panel, indistinguishable from a real failure. Such a task now fires on the next start if it came due within the last hour; older than that it is retired, marked unread, and logged. A task set with a zero-minute delay, which was retired on the spot for the same reason, runs too.
🐛 Fix: Session history for every engine, from any directory
Loading a session's history resolved the transcript against the directory Cockpit was launched from, so it failed for any session outside that directory — which is every session when the prod server starts outside a project — and it only ever looked in Claude's store, so Codex, DeepSeek, Kimi, GLM and Ollama sessions always 404'd. It now resolves by the session's own project and checks every engine's store.
🐛 Fix: Finished sessions sort above running ones
In the sidebar session dropdown, recent sessions and the mobile list, sessions that finished and are waiting for you to read now come before those still running.
📚 Docs:
/ss,/dland the Host checkThe Skills guide covers
/ssand/dl, the CLI reference documents the Host / Origin check withCOCKPIT_ALLOWED_HOSTSandCOCKPIT_DELEGATE_MAX, and a new post walks through both commands: Find any past session with /ss, hand work off with /dl.v1.0.274
2026年9月12日 · v1.0.274在 GitHub 查看 ↗✨ New: Codex streams its reply
Codex replies used to arrive as one block while every other engine typed. The cause was the transport, not the renderer:
@openai/codex-sdkspeaks onlycodex exec --experimental-json, whose event surface has no incremental text at any setting. This drops the SDK and speakscodex app-serverdirectly, which publishes genuine fragments rather than cumulative snapshots — 30 text deltas where there was 1 block, and a "processing" counter that climbs continuously instead of sitting at 0 until the very end. The client needed no changes at all; the streaming path written for Claude takes Codex verbatim.Tool calls now get a globally unique id from birth, recorded to disk, so a snapshot keyed by one is still findable after a reload. That retires five id resolvers, six per-turn counters and a shell-command fuzzy matcher which existed only because the old ids restarted at
item_0every turn — measured over the real shadow repos before removal: 1489 (session, tool id) pairs, zero duplicates.The rewrite surfaced a run of failures that all failed silently, now fixed: a failed turn reported as a clean success (failure rides on
turn/completed— there is noturn/failed), a retryable error leaving the turn hanging until aborted by hand, every sub-agent bubble's live path being dead code from a snake_case mismatch, Codex turns having lost their todo list entirely, and a sub-agent finishing tearing its parent down mid-wait_agent. Stopping a run now asks before it kills: children are interrupted first, then the parent, and the process kill is only the bound.Separately, a
thread/resumethat fails no longer restarts as a fresh thread in silence. When the ChatGPT desktop app holds a thread's writer lock, the server refuses every resume — you would get a second tab for what looks like one conversation and a model that had forgotten everything. That now lands as a system row naming both session ids and the engine's own message.✨ New: The transcript holds your reading position
Three reports, one hole. A completed turn left half a screen of blank that nothing reclaimed; a tab you had switched away from came back with that blank frozen in place; and scrolling up during a stream was undone within 50ms, parking you at the top of the turn while text kept arriving. They were the same confusion: the reserved blank counts toward
scrollHeight, so "the end of the scroller" and "the end of the transcript" differ by exactly one spacer, and the pin formula treated those two numbers as identical.The spacer is now derived from the position being held rather than owned as state — exactly enough blank for that position, never a pixel more. It evaporates as the reply grows into it, a pin no longer has to be released with a jump, and no state can hold a blank that outlives the position that justified it. Sending still scrolls your question to the top and reserves the room below it, so the answer grows into blank instead of shoving the question off screen; a long user turn (a pasted SKILL.md, routinely) clips to 16 lines with a show-more toggle.
Your place now also survives leaving. A run that finishes while you are elsewhere no longer moves the viewport of the tab you come back to, and switching projects or reloading the page restores the tab you were on along with where you were in it. The policy is a pure reducer with property tests over 200 random event sequences, plus a regression test per report.
✨ New: Maximise a chat pane
Side by side splits the panel in half, and half of a narrow window is not enough to read a long turn in. Each pane gains a maximise button that blows one column up to cover the row, with the other keeping its layout underneath, so restoring is free. Switching tabs carries the maximise along with the focus rather than dropping out of it, and closing a split cannot leave a stale "maximised" behind to resurface the next time you ask for one.
✨ New: Context controls say what they keep, and a turn can be deleted
Forking from a message now spells out the difference: Continue from here keeps all context through that turn, Only this turn keeps just that question and answer. Plan mode is labelled read-only, and a new No history context mode sends each message to the model on its own with no prior conversation — the transcript above still records everything.
A turn can also be removed outright — the question, the answer and its tool records — from the message's own row. It asks first, and refuses while the session is running.
✨ New: Quick instructions, grouped and edited as text
Quick prompts become quick instructions and gain one level of grouping; a group opens as a flyout beside the popover. Both the global and the project list are editable as plain text — one top-level entry per line, lines beginning with
-after a group name become that group's instructions:Continue Draw with text - Draw a login flow - Draw a system architectureReads fall back to the old
prompts.jsonand never write to it, so downgrading to an earlier build still finds its data.The lightning bolt was saying three different things on screen at once. It now means only "quick instructions", in both the chat and console input bars. Recent sessions takes a history icon, scheduled tasks moves off the bare clock to an alarm clock, and skills gives up the star — that mark means "favourite" elsewhere — for a joystick.
✨ New: The message index moved onto the jump capsule
The list button sat in the composer toolbar, six icons away from the prev/next steppers that do the same job at a smaller scale. It now rides the bottom jump capsule, after a divider. The dialog itself was capped at
max-w-2xl, wasting a wide display; it uses the same shell the session boards do, so width tracks the viewport and filtering no longer makes the dialog jump.Jumping also looks like what it did. Selecting a message used to light a full-width rectangle with square corners, three times the bubble's width; the flash now lands on the turn itself, scrolls instantly rather than gliding across tens of turns (which ate the highlight's whole budget before arrival), and replays when you jump to the same row twice.
🐛 Fix: One badge per session, and it survives a refresh
Every session list carried its state twice — an 8px colour dot and the round number chip beside it, two glyphs and two colour families kept in step across five files. The chip keeps the job and the dot goes; unread drops from a near-solid red to the same restrained orange wash that loading uses, so running and done now differ by motion rather than hue, and no red reads as an error where none happened. The recent-sessions and scheduled-task counts use that one style instead of their own hand-rolled pills.
The state is also right again after a reload: a running tab keeps its indicator, the tab you were on comes back active, and recent sessions no longer lose the entries worth keeping. Project rows in the sidebar now show how many sessions they hold, with the active project's badges taking priority.
🐛 Fix: Sticky headers in Scheduled Tasks
Headers pinned in the scheduled-tasks panel leaked the content scrolling underneath them.
📦 Misc
Favourite moves off the tab bar and onto the chat toolbar, where the rest of the session's actions already are.
claude-agent-sdkgoes to 0.3.268 andcodexto 0.154.0, both still exactly pinned — bumped by hand now that the scheduled bump workflow is retired.v1.0.273
2026年9月9日 · v1.0.273在 GitHub 查看 ↗✨ New: A diff opens beside the chat that produced it
Clicking "all file changes" used to swipe the Explorer into view and draw the diff over the FileBrowser — taking away the one thing you actually want next to a diff, which is the turn that wrote it. The diff is now a column inside the agent panel, sharing the row with the chat panes. Same viewer, one home instead of two.
A diff column and a second chat pane are the same right half of the panel, so they are mutually exclusive — but only by derivation, never by mutation. Opening a diff hides a split you built rather than destroying it, and closing the diff gives back the exact split that was there. On narrow screens the column can take the full panel width; the panes stay put underneath rather than being squeezed to nothing.
✨ New: Two chat panes, and they survive a reload
The agent panel splits into two panes with independent sessions. The layout is now persisted per project, so a refresh no longer silently collapses you back to a single pane — and a session closed in another browser tab leaves that pane blank instead of collapsing the split behind your back.
One composer serves both panes: the focused chat portals its input into a shared slot below them, so neither pane spends column height on it.
✨ New: Send a message to the other pane
In side-by-side mode, a message's hover row gains a fourth control between the excerpt scissors and the timestamp. It forwards that message's text to the other column and starts a run there immediately. The arrow points at the neighbour, so the button reads as a direction.
It sends the same plain text the copy button beside it would give you — no "forwarded from…" framing, since the receiving model reads framing as instructions — and it does not steal focus, because both panes are already on screen.
✨ New: The user-message list is the whole transcript now
The jump-to-message modal listed whatever the chat happened to have paged in (10 turns), so it was silently incomplete and clicking anything above that window did nothing. On one real session it showed 13 of 52 turns.
It is now served from disk: every human turn in the transcript, searchable, and selecting an unloaded row widens the chat window first so the jump actually lands. Timestamps show up for persisted rows (they were blank for everything but the current turn), and the jump highlight no longer flashes a white band around the bubble on the dark theme.
✨ New: Lean 4
.leanfiles were degrading quietly on four surfaces at once — no syntax highlighting, a generic grey file icon, line-level instead of block-level diffs, and no presence in the project graph at all. All four are fixed, including a∀file icon and theorem-level block diffs, which on a proof repo is the whole question you're asking a diff.Validated against a 60,598-file formalization repo: 434/500 exact-name matches, zero mismatches, zero overlapping spans across 76,891 lines. Symbols only — import and call edges are deliberately not scraped, because Lean resolves lemmas by elaboration and a best-effort guess would render mostly-wrong edges as fact.
🐛 Fix: Large repos stop OOM-killing the server
Opening a very large repo could take the whole process down about four minutes after the first query — post-build analytics is fire-and-forget, and on an 809K-symbol index PageRank and the tf-idf pass peaked at 10 GB over an 8 GB heap, with 82 seconds of that synchronous and the event loop dead throughout.
Three bounds fix it. The file cap goes from 15,000 to 100,000 (
git ls-filesis alphabetical, so the old budget was exhausted inside the first directory and search returned zero hits for everything later in the alphabet — which reads as a broken feature, not as a cap). Analytics is skipped above 200K symbols, reporting a distinct reason so consumers stop retrying for a result that will never arrive. And the index cache is now LRU with a 1M-symbol budget, sized by bytes rather than entry count, instead of keeping every project you ever opened resident forever.🐛 Fix: One node shape across the projectGraph API
/api/projectGraphspoke three vocabularies for the same concepts, so a parser written against one endpoint readundefinedagainst another — the right number of rows, every field empty. Search now returns the canonical node with a real line range instead of the Cmd+K palette shape,calleescalls its focal symboltargetlike its siblings, and co-edit rows agree onfilePath. Co-edit history also gains aprobability, so callers stop comparing a raw count against a ratio threshold.🌐 Site: Crawlability, a real /docs index, and consistent naming
robots.txtwas disallowing/_next/, hiding the CSS and JS Googlebot renders the page with. Every page shipped<html lang="und">, including the Chinese ones. 58 of 86 sitemap URLs carried nolastmod. The favicon and manifest 404'd. All fixed, with sitemap dates derived from real content — post dates, release dates, and the last commit that touched each docs file, recorded at author time rather than computed on CI, where a fresh clone would claim the entire site changed on every deploy./docswas a bare redirect stub, which under a static export compiles to Next's client-side error shell: 23 KB of HTML with no<h1>, no prose, and nolang— at a URL the homepage links to directly. It is now a real index page, with a card per docs page whose blurb is extracted from that page's own opening paragraph, putting all 26 docs pages one hop from the homepage.The site also now says OpenCockpit on first mention throughout — docs openings, blog titles and descriptions, homepage copy — rather than the generic noun.
📦 Misc
Chat's prose surfaces finally have a typographic contract: a bounded reading measure, leading and scale shared by chat turns, the Explorer preview, review pages, the skills modal and the diff modal. User bubbles get a brand-teal border so their edge reads as a boundary on the dark theme. The website has a lint setup of its own instead of relying on
next lint, which Next 16 removed.