SDSignal Desk

Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode

Oct 5, 2026, 3:51 PM · MarkTechPost

Image: MarkTechPost

Together AI’s new free CLI lets developers keep their favorite coding agents and swap in cheaper open models underneath, which is a real pressure point on closed-model pricing.

Why it matters

Coding agents have become a serious line item. Together’s own framing is that engineering organizations spend anywhere from tens of thousands to millions of dollars a month on closed models, often sending trivial fixes to the same premium model as full rewrites.

Together Link, released in beta under an MIT license, connects tools like Claude Code, Codex, OpenCode, Pi and the Claude and ChatGPT desktop apps to open models hosted on Together, including Kimi K3, GLM 5.3, GLM 5.3 Flash and DeepSeek V4.1 Flash. The harness stays; the model behind it changes.

From the desk

We like this for a simple reason: it separates the agent from the model. For the past year, the tool a developer loves and the model vendor that bills them have been bundled together in practice. A one-command switch that leaves normal configuration files untouched makes the model a swappable part, and swappable parts get cheaper.

The design choices are sensible. There’s no local proxy or daemon; terminal agents get a temporary per-launch configuration that goes away when the session ends, and the desktop profiles are reversible. The default auto router reads a session’s first task and picks a tier once, so prompt caching keeps working. With an Anthropic key, Claude Code sessions can route between Opus 5.5 and GLM 5.3; without one, between GLM 5.3 and its Flash variant. Each session prints its token and dollar totals on exit, and Claude Code’s status line shows estimated spend next to the equivalent Opus cost. Receipts like that are how cost claims get tested.

Those claims are Together’s own, though: over 50% savings, and 50 to 80% against all-Opus sessions. Savings only matter if quality holds on a team’s real work, and a router that decides once per session from the first task will sometimes guess wrong on a task that turns hard halfway through. I’m watching for independent reports on that failure mode.

There’s also a trust shift worth naming. Every request now flows through Together’s hosted gateway. For many teams that’s fine, but code is sensitive, and moving it to a new provider is a security and procurement decision, not just a CLI install. And mapping Claude Code’s familiar tier names onto different open models could confuse anyone who doesn’t read the docs.

Where this leads if it scales: agent harnesses become neutral shells, and model providers compete on price per finished task. That’s good for developers and for open models. It also hands a lot of leverage to whoever owns the router.

Context

Together lists Kimi K3 at $3.00 per million input tokens and $15.00 output, and GLM 5.3 at $1.40 and $4.40, with current rates on its pricing page. The tool is limited to macOS and Linux during beta, and Together notes commands, routing and the model list may change. Claude Code Router is a comparable option that offers broader provider control but runs a local gateway.

Who feels it

Developers
A low-friction way to test open models inside the agents they already use, with per-session cost visibility.
Engineering leaders
A lever on agent spend, but routing code through a new hosted gateway needs security and procurement review.
Closed-model vendors
More pressure on pricing as the agent harness and the underlying model come apart.

What to watch

  1. Independent comparisons of task success and cost between auto-routed and all-premium sessions
  2. Whether harness makers like Anthropic or OpenAI respond to third-party model swapping
  3. Windows support and how much the routing and model list change after beta

Read the original

Continue at the source.

MarkTechPost

Companies: OpenAI, Anthropic