OpenCockpit

Changelog

Release notes pulled from GitHub Releases.

  1. v1.0.282

    Oct 9, 2026 · v1.0.282

    ✨ New: Claude Haiku 5.5

    Claude Haiku 5.5 (claude-haiku-5-5) is in the Claude model picker above Haiku 4.5. It supports effort from low to max, plus ultracode and ultrathink, and defaults to medium. It costs $0.10 / $0.50 per MTok with cache reads at $0.01. Prompts over 100K tokens bill at 5x, and Token Stats only tracks totals per model, so the cost it shows for Haiku 5.5 is a lower bound.

    Engine pins: Claude Agent SDK 0.3.285 → 0.3.293 (Claude Code 2.1.293), Codex 0.159.2 → 0.161.0.

    View on GitHub ↗
  2. v1.0.281

    Oct 1, 2026 · v1.0.281

    🐛 Fix: Comments button only lights up when there is something to read

    The comments button next to the chat input used to be tinted all the time. It now uses the same muted color as the notes button until the project has at least one code comment, and picks up the tint as soon as one is added. Deleting the last comment mutes it again.

    📚 Docs: Running the agent on any model

    The engines page now opens with a direct answer to "how do I use Claude Code with another LLM". A new section compares overriding Claude Code's environment or putting a gateway in front of it with picking an engine per tab, and covers what the Built-in Agent gives up.

    🌐 Site: Two new posts and a sharper headline

    Two new blog posts: Claude Code with GLM, Kimi and DeepSeek — No Env Vars, and a side-by-side comparison of a terminal CLI workflow with the OpenCockpit workbench, including when to pick which. Page titles, the homepage lead, the social card and the footer now say "Open Claude Code GUI for any LLM".

    📦 Misc: npm description and README

    The npm description, keywords and README headings now use the same any-LLM wording as the website.

    View on GitHub ↗
  3. v1.0.280

    Sep 30, 2026 · v1.0.280

    ✨ New: Every running terminal, from the sidebar

    Terminals now outlive Cockpit restarts, which made a dev server left running in another project easy to lose track of. A Terminals row under Scheduled Tasks in the sidebar shows how many commands are live and opens a board grouped by project. Clicking a row switches to that project, swipes to Console and selects the bubble; each row can also stop its process. A stopped command stays in its bubble as finished, the same as stopping it from the bubble itself.

    ✨ New: Claude Sonnet 5.5 and GPT-6.1 Sol

    Claude Sonnet 5.5 (claude-sonnet-5-5) is in the Claude model picker, with effort up to xhigh and max, and ultracode support. It defaults to medium effort and costs $2 / $10 per MTok. Opus 5.5 stays the Claude default.

    GPT-6.1 Sol (gpt-6.1-sol) is codex's new default and is now Cockpit's default Codex model, at low reasoning.

    ✨ New: /go keeps a decision log

    /go now works from the agreed spec. It verifies each slice by running it, retries a failing slice at most three times before marking it blocked, and stops only for choices that are costly to undo. Every decision the spec did not cover is logged as one line to ~/.cockpit/skills/go/notes/<project>-<feature>.md, linked from the recap. The notes are cleaned up after 30 days without changes. /ap is gone, since its decision log now lives in /go.

    /qa, /fx and /ex now produce fixed, skimmable documents. /qa writes a lean PRD, /fx a root-cause analysis with a cited evidence chain, and /ex an answer-first analysis. Follow-up rounds show only what changed. The /cr findings index now renders as a list instead of one run-on paragraph.

    ✨ New: Baseline shifts hidden in the per-call diff list

    A turn that starts with a branch switch used to open the per-call diff viewer on a 200-file call made of other people's commits. Those leading baseline-scale calls are now hidden, as the aggregate view already did. A banner says how many were hidden and shows them again in place.

    🐛 Fix: Terminal reruns and first-frame size

    Rerunning a dev server that took longer than 200ms to exit could leave the new run marked finished with no output, because the old run's late exit was routed to it. Each run now has its own session. A new PTY also used to miss the bubble's first resize and start at 120×30, so the zsh right prompt wrapped and stray % lines appeared. It now opens at the bubble's actual size.

    This update leaves the pty-host untouched, so terminals that are already running survive it.

    🐛 Fix: Scheduled tasks under concurrent writes

    A read that landed during a scheduled-task status write saw an empty file and failed with "not valid JSON, refusing to overwrite it". This broke the task panel, and at boot it left no timers armed until the next restart. Reads now take the same file lock as writes.

    📚 Docs: Quickstart ends with /cr

    The feature-work quickstart now runs /cr before the manual review in Explorer, and the built-in skill count is down to 12.

    Engine pins: Claude Agent SDK 0.3.283 → 0.3.285, Codex 0.158.0 → 0.159.2.

    View on GitHub ↗
  4. v1.0.279

    Sep 28, 2026 · v1.0.279

    ✨ New: Terminals survive a Cockpit restart

    Terminal commands used to be children of the server, so every update or restart killed them — a dev server started in a Console bubble died with each npm i -g @surething/cockpit. They now run in a separate pty-host process (cockpit-pty in your process list), which the server talks to over a local socket. On boot the server adopts the sessions that are still running, together with the output they printed while it was down; commands that ended in the meantime come back with their real exit code. Both PTY and pipe mode are hosted.

    cockpit stop     # shuts the pty-host down too — this ends your terminals
    

    Update and restart leave the host alone. The one exception is a release that changes the host's own code: the new server then replaces the old host, which ends the terminals it was running. The host exits by itself once it has no sessions and no clients.

    Terminal output is also persisted more reliably. Finished PTY bubbles used to come back blank after a refresh, and they now keep their output. Deleting a bubble or clearing a tab removes its output file, and orphaned output files are swept at startup. \r-only progress bars in pipe mode no longer grow without bound.

    An Ollama server started from the UI survives restarts too. It used to be killed along with the server's other child processes on every stop, update or restart.

    ✨ New: An in-page folder picker

    Open Folder used to show an OS dialog launched from the background server process. Each platform broke it differently: macOS 26 would not bring it to the foreground, so paste never reached it; on Windows it opened behind the browser; and on Linux under ssh or systemd there was no display to show it on. Opened from another device, it appeared on the server's screen. The picker now lives in the project browser. It has an address bar with completion: type or paste a path, Tab/Enter goes into a directory, Cmd/Ctrl+Enter opens it, Backspace goes up and Esc returns to the project list. Git repos and hidden directories are marked.

    🐛 Fix: Wide tables no longer flip the view

    When a horizontal scroll reached the edge of a wide table, the rest of the swipe, momentum included, passed through and switched panels. A swipe now belongs to whichever pane it started over, until the wheel goes quiet. The trade-off is that a swipe over a code or diff pane with long unwrapped lines no longer switches views or dismisses the diff column. Use the top bar or the diff's ✕ instead.

    🐛 Fix: Independent-task turns no longer hide the history

    While an independent (no-history) turn was running, opening that session in another tab could replace the whole transcript with the single in-flight turn. The view then jumped to the top once the turn finished. Readers now see the stashed history and the running turn together.

    🐛 Fix: Long quick instructions stay on one line

    In the quick-instructions editor, a long instruction used to wrap and look like several separate records. Each instruction now stays on one line and scrolls horizontally, and the editor can be resized again.

    📚 Docs: /cr restructured

    The built-in /cr prompt now spells out how the review runs. The main session dispatches independent reviewers and assembles the report, but never reviews anything itself. By default there are two reviewers, one static and one dynamic; tiny diffs get a single reviewer, and hard fans out one reviewer per slice. Each finding carries a severity and an origin (introduced, activated or pre-existing), and the report has a fixed format.

    Engine pins: Claude Agent SDK 0.3.280 → 0.3.283, Codex 0.156.1 → 0.158.0.

    View on GitHub ↗
  5. v1.0.278

    Sep 23, 2026 · v1.0.278

    ✨ New: Output styles

    A global library of named output styles, picked per session from a new toolbar picker. Styles are edited in one textarea — a # Name heading per style, the body underneath taken verbatim — and stored in ~/.cockpit/output-styles.json. None is the default and injects nothing.

    The session stores only the style id; the text is resolved on every dispatch. An edit therefore reaches the next turn of every session using that style, and a scheduled task reads its session's selection at fire time rather than when it was scheduled.

    Each engine receives it in its own place: appended to Claude Code's system prompt, sent as developer instructions to Codex (reconciled on resume, since Codex ignores that parameter there), and added as a final section for the built-in agent loop. An unchanged style renders byte-identical, so Claude's prompt cache still hits after a switch back.

    This also fixes Claude sessions that had been running without Claude Code's system prompt: when no system prompt was given, the Agent SDK sent an empty custom prompt instead of the claude_code preset. The preset is now always passed.

    ✨ New: Claude Opus 5.5 and GPT-6 Sol / Luna

    Claude Opus 5.5 is the new default Claude model, at medium effort — Claude Code's own default for it, and the one model in the lineup that does not default to high. It supports xhigh, ultracode and fast mode. GPT-6 Sol is the new default Codex model at medium, with GPT-6 Luna alongside it; labels, effort sets and defaults come from the catalog Codex ships.

    Both pickers are trimmed to the current lineups. Claude drops Opus 4.8, 4.7, 4.6, Sonnet 4.6 and Fable 5, keeping Opus 5 as a fallback; Codex drops GPT-5.6-Terra and GPT-5.6-Luna, keeping GPT-5.6-Sol. A session pinned to a removed model keeps running on it, and its picker now shows the model's name (Claude Opus 4.8) rather than the raw id.

    The 200K / 1M context toggle is gone. Every model left in the picker runs a native 1M window, so it was offering a choice that changed nothing.

    Engine pins: Claude Agent SDK 0.3.278 → 0.3.280, Codex 0.155.1 → 0.156.1.

    ✨ New: Markdown preview in history diffs

    The history tab's compare mode and commit detail now offer Preview on Markdown files, opening the interactive preview on the "after" content — with comments and send-to-AI. ESC closes only the preview.

    🐛 Fix: Git status in worktrees, JSON preview in compare mode

    A commit made inside a git worktree only rewrites the branch ref in the main repository's .git, which the watcher did not look at — so the Changes tab stayed stale until something else refreshed it. It now follows the worktree's common git dir.

    In compare mode, the Readable button for JSON files did nothing until you switched to the status tab. It now opens the preview directly.

    🐛 Fix: Command tags no longer change with the UI language

    Resolved command tags read [主会话·qa] next to [subagent·cr] for a Chinese UI, and scheduled tasks fell back to English regardless. The tag is model-facing metadata, not UI copy, so it is now main / subagent in every language.

    🌐 Site: Long-term memory is just a directory

    A new blog post, in English and Chinese, on why Bots keep memory as a plain directory the model explores like a codebase — with BOT.md as the index — instead of a vector database. The /try demo, which failed to start after the cock alias removal, is fixed too, as is an install snippet on the site that still used the old name.

    View on GitHub ↗
  6. v1.0.277

    Sep 21, 2026 · v1.0.277

    ✨ New: Read a whole turn as one diff

    The chat diff viewer only ever answered "what did this tool call do". Reading a turn meant clicking through every call and summing the overlaps by eye. An aggregate toggle now collapses the range into a single diff, built over the shadow snapshot repo from the first call's parent to the newest one.

    It is a net diff, and that is the point: a file edited five times appears once, and a file created then deleted inside the range does not appear at all. The count therefore drops — 41 per-call entries can be 12 distinct files — so the header labels the aggregate rather than letting that read as lost data.

    A turn that opens with git checkout -b x origin/main used to start the range on a call that rewrote 200 files and 17k lines of other people's commits. Nothing in a snapshot says "this was a branch switch", so size stands in for intent: oversized leading calls are left out of the starting point. That is a guess, so it is never silent — the meta bar states how many calls were dropped and toggles them back in.

    Two interaction changes ride along, both about closing the viewer. The header controls move to the left, macOS-window style, since the pointer already lives over the left half of the diff and a ✕ in the far corner charged a full-width trip for the most common action. And the diff column can now be swiped right to dismiss: the gesture draws itself, the column trailing the swipe at half its travel before flying out or springing back.

    ✨ New: Branch compare counts work you have not committed yet

    The history tab's branch compare ran git diff <base> HEAD, so a branch whose work was still in the working tree rendered "no changed files" while a pile of fresh edits sat on disk. Between that view and the status tab, nothing answered "what did this branch change in total" — the question an agent-driven session asks most.

    The HEAD pill in the compare header is now a two-segment scope switch:

    • worktree — committed + staged + unstaged + untracked (the new default)
    • head — committed work only, the GitHub "Files changed" diff

    Untracked files are merged in from git status, since git diff never reports them. Both modes now anchor the old side at the merge base; the previous two-dot spelling reported commits made on the base branch since branching as if this branch had reverted them. A project opened at a subdirectory of its repo also read working-tree paths against the wrong root and returned an empty diff.

    ✨ New: A Bot session says so, wherever it is listed

    A Bot session used to be indistinguishable from one a person opened. The mark now rides on chrome that already exists rather than adding a glyph: the number chip keeps its number and changes shape — a robot head, masked over the same status wash — and the running line gets a badge. Detection reads the @name out of the dispatch line once, server-side, and it travels with the payload that already answers which engine.

    An @bot line sent from a session that is already sitting in that Bot's directory now runs in place instead of spawning a child. Delegating exists to give a Bot a session whose files are its own; a session that already has that bought a cold start and a second transcript and nothing else. Matching is exact after realpath, so a subdirectory of the Bot does not match and a repo root containing bots/ cannot swallow every @bot line in the project. A delegation with no explicit engine also inherits the engine of the session that asked for it.

    🐛 Fix: Ollama runs the model your machine actually has

    Ollama ids are whatever the machine pulled, so any id baked into the source is a guess that ollama pull / rm invalidates — and one delegation could produce three separate 404s from Cockpit rather than from the model. Model resolution now happens against the machine: the request's own model (a bare gpt-oss completed to its one installed tag), else the model this session already ran on, else the model the last run used, else the catalog's first entry. An empty or unreachable catalog is now an error naming the server URL and the fix, raised before the run starts rather than landing in a transcript you then throw away.

    A new ollama chat opens on the model last used instead of on "Select model", and the picker and the engine read the same resolver, so they cannot answer the question differently.

    🐛 Fix: A scheduled task follows its fresh session immediately

    When a task's resume target is gone it starts a fresh session, and that session's id used to be written back only after the run finished, and only on the success path. So for the whole turn the task pointed at a session that no longer exists — the board opened an empty transcript while the real one streamed under another id — and a run that timed out left the dead id on disk, which is self-reinforcing: the next round starts yet another from-scratch session, and a from-scratch run is the one most likely to fail again. The task now binds on the id the engine announces seconds in. The per-run deadline also rises to 60 minutes; it is a runaway guard, not a latency budget.

    🐛 Fix: Quick instructions can span several lines

    The outline editor's record separator is the line break, so a multi-line instruction had no way to spell itself — and the round trip was already lossy, silently tearing such an instruction in two on the next save. Real newlines are now encoded as a literal \n inside the textarea and decoded on the way out, so storage, the API and the send path keep carrying real newlines.

    📦 Misc: The cock alias is gone

    cock is an English profanity, which makes the short alias awkward in shell history, CI logs, documentation and talks. Cockpit ships a single name now:

    cockpit          # production server
    cockpit-dev      # dev server
    

    If you had cock in a script or an alias, switch it to cockpit — same 237-line implementation, now living in bin/cockpit.mjs. The lockfile carried a stale entry pointing at the deleted file, which would have left npm ci with a dangling symlink.

    Engine pins move with this release too: Claude Agent SDK 0.3.268 → 0.3.278, Codex 0.154.0 → 0.155.1.

    View on GitHub ↗
  7. v1.0.276

    Sep 18, 2026 · v1.0.276

    ✨ New: Bots — a folder you address as @name

    A Bot is a plain directory with a BOT.md at its root. Register it, then start a line with @name: that message runs as its own session, carrying the Bot's files as long-term context.

    @reviewer look at the auth changes on this branch
    

    Cockpit does three things and no more — resolve the name, rewrite the message, expose the delegate and status endpoints. It never reads, writes or locks a Bot directory, so a Bot stays an ordinary folder you can edit by hand, keep in git, or hand to someone else.

    The command surface is now split by prefix: /verb runs in the main session, /@verb is delegated, @name is a Bot. Create one with /bot. Writes are explicit and locked — a Bot's files change only when you ask it to remember, update, correct or forget, under <bot>/.locks/write, whose owner file records the run so a stuck lock is decided by asking whether that run is still alive rather than by a timer.

    The full design, including the failure mode each rule exists to prevent, is in docs/BOTS.md and at Bots.

    ✨ New: Share a Bot with @name export to <path>

    A Bot's working directory is the wrong thing to share. It holds the write lock, review reports, your own memory, and paths that exist on exactly one machine — hand it over and all four go with it, usually into a public repository where the mistake cannot be taken back.

    @reviewer export to ~/share/reviewer
    

    Export builds a new directory instead: BOT.md rewritten, identity/ and skills/ carried over, everything else recreated as empty scaffolding. It is a whitelist rather than a blacklist, so a Bot that grows a notes/ next month does not quietly start shipping it. The report ends with the Skill names the copy references, which is the whole installation requirement for whoever receives it. What happens to that directory afterwards is your business — Cockpit neither publishes nor registers it.

    A Bot's Skills table now names its tools instead of locating them: a registered skill's name, a Bot-relative path for one the Bot grew itself, an absolute path only when it is neither — flagged as machine-local when it is written. That is what makes a shared table mean anything. The recipient registers the skill under the same name, their own copy in their own location, and the row resolves for them as it did for you. No path could do that, because the only paths two machines share are the ones neither of them chose.

    ✨ New: Built-in Bots live in ~/.cockpit/bots

    A built-in Bot used to run from the install root — root-owned under npm i -g, replaced wholesale on every upgrade, shared by every COCKPIT_HOME on the machine. The panel sent you there anyway, so an edit was either refused outright or silently deleted by the next upgrade.

    The shipped directory is now a seed. Each built-in is installed into ~/.cockpit/bots/<name>, and that copy is the Bot: the panel card, the folder button and @name dispatch all follow the same path to a directory you can open and edit. Install tracking is per file, so editing the persona keeps your edit and still takes later BOT.md fixes — where before, one edit froze the whole Bot at the version you touched it. Deleting the folder resets it.

    cockpit-helper, which answers questions about Cockpit by reading opencockpit.dev live, now keeps memory as a result — of the person, never of the site. Which install you run, how you want answers, corrections you made, what you asked it to follow up on. Not a remembered route through the docs, which is how it would stop reading the site and start guessing at it.

    🐛 Fix: A malformed registry is no longer overwritten with emptiness

    bot.json, skills.json and scheduled-tasks.json are files you are invited to edit by hand, so a stray comma is a realistic state rather than a hypothetical one — and it read as "empty", after which the next write persisted that emptiness. A corrupt file now fails the write instead of replacing your data behind a success toast.

    The scheduler's boot path is the deliberate exception: it catches, logs that no task will fire, and refuses to save while in that state. Propagating the throw there would have traded silent data loss for a Cockpit that will not start.

    🐛 Fix: Cockpit's own PORT no longer leaks into your projects

    The launcher exports its listening port as the generic PORT, which Next, Vite and most dev servers read as their port. It is now stripped from every process spawned into a project. COCKPIT_PORT stays, for CLI bridges that want it, and a custom port still survives a restart or update.

    Each engine also exports COCKPIT_CWD now — the directory the session was started with. Everything keyed on cwd (delegation, the Bot write-lock owner, status lookups) previously had to work it out, and a session running in a subdirectory worked it out wrong: it saw the repository root's .git, called that "the project", and wrote records nobody could look up. pwd is not the answer either, since it follows any cd the turn has made.

    ✨ Chat and Git polish

    The branch compare header now carries summed +additions / −deletions on its right, next to "N files changed", styled like the per-file counts. The <engine> running <elapsed> line gets an orange shimmer sweeping across it, matching the running spinner — @supports-guarded, and off under prefers-reduced-motion. Notes now sit before Skills in the sidebar, and the Git change-classification badges line up.

    🌐 Site: SEO metadata and breadcrumbs

    Per-page canonical and OpenGraph metadata for docs and blog, breadcrumb markup, a corrected robots.txt, and a validation pass in the site build. Docs <lastmod> dates were also re-derived from git — 22 of them had drifted, which matters because a sitemap whose dates are learned to be wrong gets trusted less.

    View on GitHub ↗
  8. v1.0.275

    Sep 14, 2026 · v1.0.275

    ✨ New: /ss finds a past session from one sentence

    You remember what a conversation was about, not which project, engine or week it happened in. /ss takes that half-memory and finds the session:

    /ss the session where we worked out CSRF on the local API
    

    The agent does not search your sentence verbatim. It expands it into the words that would literally appear in that conversation — both languages for technical topics (跨站 / CSRF), synonyms, the terms the assistant would have used — searches every session Cockpit can read across all projects, all engines (Claude, Codex, DeepSeek, Kimi, GLM, Ollama) and all dates, reads the snippets, and replies with 1–3 candidates. Each carries a session link; clicking it switches to that project and opens the session in the Agent panel.

    Behind it, Cockpit keeps a text-only copy of each session's prompts and replies under <data-dir>/search-corpus, updated incrementally, and searches it with ripgrep, so two-character Chinese words like 快照 match. The first search on a machine takes a few seconds while that copy is built.

    ✨ New: /dl hands work to another session without waiting

    Halfway through a task you notice work that belongs somewhere else — a flaky test in another repo, a job better suited to Codex. /dl starts it there and returns at once:

    /dl have codex fix the flaky date test in the api project
    

    The agent writes a self-contained brief (the child sees none of your conversation), starts a new session in the target directory on the chosen engine, and repeats the receipt — engine, directory, link — in its reply before carrying on. The child is an ordinary session: open the link to watch it or take over. Later, ask how it went; the agent finds the receipt, in this conversation or through /ss, and reports running, done, failed or incomplete along with the child's last reply.

    Nothing is stored server-side — status is read from the child engine's own transcript. At most 4 delegated sessions run at once (COCKPIT_DELEGATE_MAX); past that the request is rejected rather than queued.

    ✨ New: A Host / Origin check on every request

    A web page open in your browser could previously reach a local Cockpit through DNS rebinding or a cross-site request. Every HTTP request and WebSocket upgrade now passes a check first: a request arriving over loopback must address Cockpit by a local host name, and a state-changing request or WebSocket upgrade that carries an Origin must be same-origin. Rejections are 403 Forbidden. curl, the CLI and skills send no Origin and are unaffected.

    If you reach Cockpit through a tunnel (ngrok, cloudflared, …), this changes behaviour. The tunnel connects over loopback with its public hostname in Host, so it is now refused unless you allow that name or run in token mode:

    COCKPIT_ALLOWED_HOSTS=my-box.ngrok.app cockpit     # no token: anyone with the URL gets in
    cockpit --token my-secret-value                    # token mode: the tunnel must forward X-Forwarded-For
    

    Keep the tunnel's original Host header; if it is rewritten to localhost, POSTs and WebSockets get 403.

    🐛 Fix: A one-time task that came due during a restart now runs

    Scheduled timers die with the process, and a one-time task whose moment passed while Cockpit was down was marked completed without ever running — and then showed as failed in the panel, indistinguishable from a real failure. Such a task now fires on the next start if it came due within the last hour; older than that it is retired, marked unread, and logged. A task set with a zero-minute delay, which was retired on the spot for the same reason, runs too.

    🐛 Fix: Session history for every engine, from any directory

    Loading a session's history resolved the transcript against the directory Cockpit was launched from, so it failed for any session outside that directory — which is every session when the prod server starts outside a project — and it only ever looked in Claude's store, so Codex, DeepSeek, Kimi, GLM and Ollama sessions always 404'd. It now resolves by the session's own project and checks every engine's store.

    🐛 Fix: Finished sessions sort above running ones

    In the sidebar session dropdown, recent sessions and the mobile list, sessions that finished and are waiting for you to read now come before those still running.

    📚 Docs: /ss, /dl and the Host check

    The Skills guide covers /ss and /dl, the CLI reference documents the Host / Origin check with COCKPIT_ALLOWED_HOSTS and COCKPIT_DELEGATE_MAX, and a new post walks through both commands: Find any past session with /ss, hand work off with /dl.

    View on GitHub ↗
  9. v1.0.274

    Sep 12, 2026 · v1.0.274

    ✨ New: Codex streams its reply

    Codex replies used to arrive as one block while every other engine typed. The cause was the transport, not the renderer: @openai/codex-sdk speaks only codex exec --experimental-json, whose event surface has no incremental text at any setting. This drops the SDK and speaks codex app-server directly, which publishes genuine fragments rather than cumulative snapshots — 30 text deltas where there was 1 block, and a "processing" counter that climbs continuously instead of sitting at 0 until the very end. The client needed no changes at all; the streaming path written for Claude takes Codex verbatim.

    Tool calls now get a globally unique id from birth, recorded to disk, so a snapshot keyed by one is still findable after a reload. That retires five id resolvers, six per-turn counters and a shell-command fuzzy matcher which existed only because the old ids restarted at item_0 every turn — measured over the real shadow repos before removal: 1489 (session, tool id) pairs, zero duplicates.

    The rewrite surfaced a run of failures that all failed silently, now fixed: a failed turn reported as a clean success (failure rides on turn/completed — there is no turn/failed), a retryable error leaving the turn hanging until aborted by hand, every sub-agent bubble's live path being dead code from a snake_case mismatch, Codex turns having lost their todo list entirely, and a sub-agent finishing tearing its parent down mid-wait_agent. Stopping a run now asks before it kills: children are interrupted first, then the parent, and the process kill is only the bound.

    Separately, a thread/resume that fails no longer restarts as a fresh thread in silence. When the ChatGPT desktop app holds a thread's writer lock, the server refuses every resume — you would get a second tab for what looks like one conversation and a model that had forgotten everything. That now lands as a system row naming both session ids and the engine's own message.

    ✨ New: The transcript holds your reading position

    Three reports, one hole. A completed turn left half a screen of blank that nothing reclaimed; a tab you had switched away from came back with that blank frozen in place; and scrolling up during a stream was undone within 50ms, parking you at the top of the turn while text kept arriving. They were the same confusion: the reserved blank counts toward scrollHeight, so "the end of the scroller" and "the end of the transcript" differ by exactly one spacer, and the pin formula treated those two numbers as identical.

    The spacer is now derived from the position being held rather than owned as state — exactly enough blank for that position, never a pixel more. It evaporates as the reply grows into it, a pin no longer has to be released with a jump, and no state can hold a blank that outlives the position that justified it. Sending still scrolls your question to the top and reserves the room below it, so the answer grows into blank instead of shoving the question off screen; a long user turn (a pasted SKILL.md, routinely) clips to 16 lines with a show-more toggle.

    Your place now also survives leaving. A run that finishes while you are elsewhere no longer moves the viewport of the tab you come back to, and switching projects or reloading the page restores the tab you were on along with where you were in it. The policy is a pure reducer with property tests over 200 random event sequences, plus a regression test per report.

    ✨ New: Maximise a chat pane

    Side by side splits the panel in half, and half of a narrow window is not enough to read a long turn in. Each pane gains a maximise button that blows one column up to cover the row, with the other keeping its layout underneath, so restoring is free. Switching tabs carries the maximise along with the focus rather than dropping out of it, and closing a split cannot leave a stale "maximised" behind to resurface the next time you ask for one.

    ✨ New: Context controls say what they keep, and a turn can be deleted

    Forking from a message now spells out the difference: Continue from here keeps all context through that turn, Only this turn keeps just that question and answer. Plan mode is labelled read-only, and a new No history context mode sends each message to the model on its own with no prior conversation — the transcript above still records everything.

    A turn can also be removed outright — the question, the answer and its tool records — from the message's own row. It asks first, and refuses while the session is running.

    ✨ New: Quick instructions, grouped and edited as text

    Quick prompts become quick instructions and gain one level of grouping; a group opens as a flyout beside the popover. Both the global and the project list are editable as plain text — one top-level entry per line, lines beginning with - after a group name become that group's instructions:

    Continue
    Draw with text
    - Draw a login flow
    - Draw a system architecture
    

    Reads fall back to the old prompts.json and never write to it, so downgrading to an earlier build still finds its data.

    The lightning bolt was saying three different things on screen at once. It now means only "quick instructions", in both the chat and console input bars. Recent sessions takes a history icon, scheduled tasks moves off the bare clock to an alarm clock, and skills gives up the star — that mark means "favourite" elsewhere — for a joystick.

    ✨ New: The message index moved onto the jump capsule

    The list button sat in the composer toolbar, six icons away from the prev/next steppers that do the same job at a smaller scale. It now rides the bottom jump capsule, after a divider. The dialog itself was capped at max-w-2xl, wasting a wide display; it uses the same shell the session boards do, so width tracks the viewport and filtering no longer makes the dialog jump.

    Jumping also looks like what it did. Selecting a message used to light a full-width rectangle with square corners, three times the bubble's width; the flash now lands on the turn itself, scrolls instantly rather than gliding across tens of turns (which ate the highlight's whole budget before arrival), and replays when you jump to the same row twice.

    🐛 Fix: One badge per session, and it survives a refresh

    Every session list carried its state twice — an 8px colour dot and the round number chip beside it, two glyphs and two colour families kept in step across five files. The chip keeps the job and the dot goes; unread drops from a near-solid red to the same restrained orange wash that loading uses, so running and done now differ by motion rather than hue, and no red reads as an error where none happened. The recent-sessions and scheduled-task counts use that one style instead of their own hand-rolled pills.

    The state is also right again after a reload: a running tab keeps its indicator, the tab you were on comes back active, and recent sessions no longer lose the entries worth keeping. Project rows in the sidebar now show how many sessions they hold, with the active project's badges taking priority.

    🐛 Fix: Sticky headers in Scheduled Tasks

    Headers pinned in the scheduled-tasks panel leaked the content scrolling underneath them.

    📦 Misc

    Favourite moves off the tab bar and onto the chat toolbar, where the rest of the session's actions already are. claude-agent-sdk goes to 0.3.268 and codex to 0.154.0, both still exactly pinned — bumped by hand now that the scheduled bump workflow is retired.

    View on GitHub ↗
  10. v1.0.273

    Sep 9, 2026 · v1.0.273

    ✨ New: A diff opens beside the chat that produced it

    Clicking "all file changes" used to swipe the Explorer into view and draw the diff over the FileBrowser — taking away the one thing you actually want next to a diff, which is the turn that wrote it. The diff is now a column inside the agent panel, sharing the row with the chat panes. Same viewer, one home instead of two.

    A diff column and a second chat pane are the same right half of the panel, so they are mutually exclusive — but only by derivation, never by mutation. Opening a diff hides a split you built rather than destroying it, and closing the diff gives back the exact split that was there. On narrow screens the column can take the full panel width; the panes stay put underneath rather than being squeezed to nothing.

    ✨ New: Two chat panes, and they survive a reload

    The agent panel splits into two panes with independent sessions. The layout is now persisted per project, so a refresh no longer silently collapses you back to a single pane — and a session closed in another browser tab leaves that pane blank instead of collapsing the split behind your back.

    One composer serves both panes: the focused chat portals its input into a shared slot below them, so neither pane spends column height on it.

    ✨ New: Send a message to the other pane

    In side-by-side mode, a message's hover row gains a fourth control between the excerpt scissors and the timestamp. It forwards that message's text to the other column and starts a run there immediately. The arrow points at the neighbour, so the button reads as a direction.

    It sends the same plain text the copy button beside it would give you — no "forwarded from…" framing, since the receiving model reads framing as instructions — and it does not steal focus, because both panes are already on screen.

    ✨ New: The user-message list is the whole transcript now

    The jump-to-message modal listed whatever the chat happened to have paged in (10 turns), so it was silently incomplete and clicking anything above that window did nothing. On one real session it showed 13 of 52 turns.

    It is now served from disk: every human turn in the transcript, searchable, and selecting an unloaded row widens the chat window first so the jump actually lands. Timestamps show up for persisted rows (they were blank for everything but the current turn), and the jump highlight no longer flashes a white band around the bubble on the dark theme.

    ✨ New: Lean 4

    .lean files were degrading quietly on four surfaces at once — no syntax highlighting, a generic grey file icon, line-level instead of block-level diffs, and no presence in the project graph at all. All four are fixed, including a ∀ file icon and theorem-level block diffs, which on a proof repo is the whole question you're asking a diff.

    Validated against a 60,598-file formalization repo: 434/500 exact-name matches, zero mismatches, zero overlapping spans across 76,891 lines. Symbols only — import and call edges are deliberately not scraped, because Lean resolves lemmas by elaboration and a best-effort guess would render mostly-wrong edges as fact.

    🐛 Fix: Large repos stop OOM-killing the server

    Opening a very large repo could take the whole process down about four minutes after the first query — post-build analytics is fire-and-forget, and on an 809K-symbol index PageRank and the tf-idf pass peaked at 10 GB over an 8 GB heap, with 82 seconds of that synchronous and the event loop dead throughout.

    Three bounds fix it. The file cap goes from 15,000 to 100,000 (git ls-files is alphabetical, so the old budget was exhausted inside the first directory and search returned zero hits for everything later in the alphabet — which reads as a broken feature, not as a cap). Analytics is skipped above 200K symbols, reporting a distinct reason so consumers stop retrying for a result that will never arrive. And the index cache is now LRU with a 1M-symbol budget, sized by bytes rather than entry count, instead of keeping every project you ever opened resident forever.

    🐛 Fix: One node shape across the projectGraph API

    /api/projectGraph spoke three vocabularies for the same concepts, so a parser written against one endpoint read undefined against another — the right number of rows, every field empty. Search now returns the canonical node with a real line range instead of the Cmd+K palette shape, callees calls its focal symbol target like its siblings, and co-edit rows agree on filePath. Co-edit history also gains a probability, so callers stop comparing a raw count against a ratio threshold.

    🌐 Site: Crawlability, a real /docs index, and consistent naming

    robots.txt was disallowing /_next/, hiding the CSS and JS Googlebot renders the page with. Every page shipped <html lang="und">, including the Chinese ones. 58 of 86 sitemap URLs carried no lastmod. The favicon and manifest 404'd. All fixed, with sitemap dates derived from real content — post dates, release dates, and the last commit that touched each docs file, recorded at author time rather than computed on CI, where a fresh clone would claim the entire site changed on every deploy.

    /docs was a bare redirect stub, which under a static export compiles to Next's client-side error shell: 23 KB of HTML with no <h1>, no prose, and no lang — at a URL the homepage links to directly. It is now a real index page, with a card per docs page whose blurb is extracted from that page's own opening paragraph, putting all 26 docs pages one hop from the homepage.

    The site also now says OpenCockpit on first mention throughout — docs openings, blog titles and descriptions, homepage copy — rather than the generic noun.

    📦 Misc

    Chat's prose surfaces finally have a typographic contract: a bounded reading measure, leading and scale shared by chat turns, the Explorer preview, review pages, the skills modal and the diff modal. User bubbles get a brand-teal border so their edge reads as a boundary on the dark theme. The website has a lint setup of its own instead of relying on next lint, which Next 16 removed.

    View on GitHub ↗