SDSignal Desk

These Russian Mathematicians Taught AI Models How to Talk to Each Other Without Using Words

Sep 2, 2026, 11:20 AM · WIRED

Image: WIRED

WIRED's Mostik profile is a lab demo: weights talking to weights, a 753B–4B hybrid at one-twentieth the cost, and a Fields Medalist as chief scientist.

Why it matters

Will Knight reports on Mostik, a Russian startup (the name means 'bridge') whose approach lets models interact using values in their weights rather than text. Capabilities of a larger model can be fed to a smaller one more efficiently. The company used the method on a system that Knight says has rocketed to the top of ARC-AGI 3; they would not say more because they want to win the contest. As a demonstration they bridged the largest GLM-5.2 (753 billion parameters) and a 4-billion-parameter Qwen-3.5 that can run on a phone. The hybrid costs one-twentieth of full GLM, with performance exactly halfway between the two.

CEO Sasha Malysheva argues ensembles beat individuals and that the future is not a monolithic scaled model. Typical ensembling feeds text from one model into another; Mostik skips the text. Vladimir Arustamian of Lovable said he would have guessed this was years out. Karl Tuyls, formerly of DeepMind, called it a no-brainer for efficiency if you can pair frontier models with domain-specific ones. Chief scientist Stanislav Smirnov, a 2010 Fields Medalist at Geneva, said there is not yet an appropriate mathematical language for common ground between models. The write-up is an edition of Knight's AI Lab newsletter.

The Signal Desk read

This is a reporter-access piece, not a paper. ARC-AGI 3 'top' without a documented submission is a claim to verify on the leaderboard, and the team is withholding the recipe on purpose. The GLM–Qwen hybrid is the only numbered result: 1/20th the cost, midpoint quality. That is a striking interpolation if it holds, and a very specific pair of Chinese open weights.

Skipping tokens between models is the idea. If it generalizes, open-weight catalogs get more valuable — you compose specialists instead of renting one frontier API. If it is a brittle, hand-tuned bridge between two named checkpoints, it is a paper waiting to happen, not a platform. Smirnov saying the math language does not exist yet is the honest sentence. They are engineering around a missing theory.

Signal Desk's read: treat Mostik as an existence proof under Knight's eyes, not as the end of scaling. Malysheva's anti-monolith line is what you would say if your product is a combiner. It can still be true. Closed labs will not federate weights with strangers; open models might. The geopolitical aftertaste — Russian mathematicians, Chinese open weights, a Swiss Fields Medalist, an American magazine — is inescapable and not the technical story. The technical story is whether 'halfway to 753B at 4B-plus-bridge cost' survives contact with tasks that are not the demo.

Do not write 'machine telepathy' into a procurement doc. Knight used it as color.

Context

Ensembling models by passing text is old, slow, and expensive. Sharing activations or weight-space features is an active research bet. ARC-AGI 3 is a hard public contest, which is why a 'we're winning but we won't explain' line travels.

Who feels it

Open-weight users
If a 4B-plus-bridge can borrow a 753B, on-device plus a specialist becomes a real architecture. Demand a third pair, not GLM–Qwen only.
Frontier labs
A combiner that works is a reason smaller open models stay relevant. It is not a reason to stop scaling, yet.
Researchers
Smirnov's missing language is the paper. Without it this stays a lab trick.

What to watch

  1. An ARC-AGI 3 listing that names Mostik, with a score that stays on top.
  2. A public method paper, or continued secrecy until the contest ends.
  3. A hybrid that is not GLM-5.2 plus Qwen-3.5.

Read the original

Continue at the source.

WIRED