Build the workbench that survives a model switch.
Most people will use more than one model. A local model may be right for one task; a hosted model, coding plan, or different agent engine may be better for the next. Cloaky keeps the project, thread, permissions, review, and cost in view while the intelligence changes.
Keep the harness useful.
An agent is not just a chat box. It reads a project, makes a plan, runs tools, asks for permission, changes files, hands work to subagents, and leaves a diff someone has to review. Cloaky holds that loop together and shows the engine, provider, permission, file change, and cost at the point they matter.
The model can change. The work should not disappear.
Use local inference when the boundary matters, a hosted route when it fits, or another engine when it has the capability you need. Keep the same habits: choose the route, watch the action, review the result, and know what the session cost.
The rules that keep the product honest.
The route is part of the result
Every session names the engine, provider, and privacy tier. A model choice changes who handles the work and under which terms.
The agent can stop
Shell commands and other consequential actions can pause with the exact command and scope in view. The important button is sometimes “not yet.”
Review is another turn
The diff belongs in the conversation. Comment on a line, send the request back, or rewind the conversation, files, or both.
Receipts beat vibes
Session cost, tool activity, files changed, tests, open issues, and subagent work should survive “done.”
Privacy claims name their owner
Local inference, Venice private, anonymized routes, and direct providers are different paths. We say which one, what it promises, and where the edge is.
The order of the work matters.
Not a calendar of promises. A view of what we need to earn next. The changelog records what is real.
Make the current loop dependable
Finish the rough edges around what already exists: real sessions, local and hosted routes, engine choice, active work, approvals, review, security findings, artifact preview, skills, workflows, schedules, and exact receipts. This is the part users can touch today.
Shipped surface · beta work continuesMove the desktop experience to Rust + GPUI
The move is about feel, not a new logo on an architecture diagram. We want faster startup, less drag in long sessions, and a surface that stays responsive with transcripts, review panes, model pickers, and several active runs open. We will show measured before-and-after results.
Next foundation · performance to be measuredFind the real privacy boundary
When a product handles code and intelligence, “private” is not a color or a badge. We are tracing each path: session turns, headlines, assists, subagents, tools, engine vendors, and account systems. The output should be better controls, clearer docs, and fewer claims we cannot prove.
Research underway · evidence before claimsCarry the work forward
Once the foundation is right, the question changes from “which model is best?” to “how does my work survive the next one?” We want the project, decisions, review trail, and receipts to remain useful as engines, providers, and devices change. That is a direction, not a promised feature list.
Further out · continuation over noveltyTurn the research into product defaults
Privacy work only counts if it changes the experience: clearer route choices, safer boundaries, useful warnings, and a better answer when someone asks where a piece of work went. There is no final checkbox here. The workbench has to keep earning trust as the intelligence around it changes.
Open-ended by design · product over postureThe roadmap does not end with the foundation. Every new engine or surface has to improve the same loop: choose the intelligence, see what it can touch, let it act, inspect what it changed, and keep the record.
Build it, use it, then tell the truth about it.
Cloaky is young. That is the operating condition, not a disclaimer. We ship a surface, use it on real work, measure what it did, write down the limit, and decide what the next build has earned. The labels stay plain: shipped, being tested, under research, or planned.
