Weight-level model bridging could shrink AI inference costs

By Zak and the True Work Office team | Published: 28 September 2026 | Category: blog | 3 min read

Weight-level model bridging could shrink AI inference costs

Key points
  • Mostik's method transfers capability between AI models through numerical values in their weights, bypassing text-based communication.
  • A hybrid pairing GLM-5.2 and a mobile-ready Qwen-3.5 variant reportedly costs one-twentieth of running the full large model while delivering intermediate performance.
  • The approach was applied to an ARC-AGI 3 benchmark entry, though technical details were withheld during competition.
  • A weight-level bridge could complicate model auditing, since interactions bypass the readable text trace that traditional ensembling produces.

A Russian startup called Mostik says it has found a way to let artificial intelligence models hand off capability to one another through the raw numbers in their neural-network weights, rather than by translating knowledge into text first. If the claim holds up, it is a materially different way to combine models, and one that could narrow the gap between large frontier systems and the smaller, cheaper models that run on phones or local hardware.

The demonstration most closely described in Wired’s reporting on the weight-based communication method involved bridging China’s 753-billion-parameter GLM-5.2 with a 4-billion-parameter Qwen-3.5 variant small enough for mobile devices. The hybrid result, according to the account, cost roughly one-twentieth as much as running the full GLM model while landing between the two components in measured performance. The approach also featured in an ARC-AGI 3 benchmark entry, though Mostik withheld technical detail while competing. That omission is understandable in a contest setting but means the method has not yet been subjected to sustained outside scrutiny. CEO Sasha Malysheva, who developed the technique, argues that composing many models may matter more than simply scaling a single one.

The practical stakes are less about funding and more about governance and evidence. If models can share capability at the weight level without routing every exchange through generated text, the cost of combining specialised systems drops. That is useful for organisations that cannot afford full frontier-scale inference, including universities running student-facing tools on modest hardware. It also complicates the auditing picture. Text-based interaction between models leaves a readable trace; a weight-level bridge does not, at least not without tooling that barely exists yet. As Fields Medallist Stanislav Smirnov notes in the same account, a shared mathematical language for model-to-model interaction remains underdeveloped, which is a polite way of saying the field does not yet have good answers for what happens when models talk to each other in a register humans cannot easily inspect.

The practical test is whether the approach generalises beyond a narrow benchmark or a single showcase pair. A hybrid that runs a quarter of the cost but only matches mid-tier performance on one task is interesting; one that reliably lifts smaller models across diverse domains without degrading safety or accuracy would be significant. Until independent replication and clearer documentation arrive, the claim sits in the useful-but-unproven category.

Frequently asked questions

What is weight-level model bridging?

Mostik’s method transfers capability between AI models through the raw numerical weights in their neural networks, rather than translating knowledge into generated text first. This lets two models share specialised ability without the overhead of converting everything to language and back.

How much cheaper could this make AI inference?

In one demonstration, pairing a 753-billion-parameter GLM-5.2 with a 4-billion-parameter Qwen-3.5 variant reportedly cost roughly one-twentieth as much as running the full large model alone, while delivering performance between the two components.

Why does weight-level bridging complicate model auditing?

Text-based interaction between models leaves a readable trace that auditors can inspect. A weight-level bridge bypasses that trace, meaning the exchange of capability happens in a register humans cannot easily scrutinise, at least not without tooling that barely exists yet.

Is this approach proven?

Not yet. The demonstration showed promising results on one benchmark pair, but Mostik withheld technical details while competing in the ARC-AGI 3 contest. Independent replication and broader documentation across diverse domains are needed before the claim moves beyond useful-but-unproven.

โ† Back to Blog