Together AI leads token share on OpenRouter; MiniMax-M3 delivers million-token context at one-tenth the cost

Together AI tops token share on OpenRouter among leading open-source coding models, according to data published today. Three models split the lead: Z.ai's GLM 5.3 Flash at 29.2%, DeepSeek's V4.1 Flash at 25.6%, and Moonshot's Kimi K3 at 18.9%. The figures reflect actual developer usage running code agents on open models, not downloads or synthetic benchmarks.
At the same time, MiniMax is launching M3, a multimodal foundation model that accepts text, image, and video input with text output and a million-token context window. The core technology is MiniMax Sparse Attention (MSA), which swaps full attention for KV-block selection. That cuts per-token compute cost in long context to 5% of the previous generation at a million tokens, with significantly faster prefill and decode while preserving quality on most tasks.
The model was trained natively multimodal on interleaved data and fine-tuned for extensive multi-turn dialogue using an interactive user-simulation framework that mimics production workloads. The goal is to steer the model toward sustained, multi-step tasks — agents that write code, invoke tools, and maintain context over time — rather than single-shot prompt completion.
For developers running code agents on open models, the OpenRouter data signals an ecosystem converging around three dominant Chinese players, while MiniMax offers an alternative with a massive context window at a cost that makes long workloads viable without breaking the budget. The per-token cost drop from MSA rewrites the economics of long-context inference, letting a full coding-session history stay resident in memory.
The emerging picture is a market where real-world performance, not benchmark promises, determines usage share, and where sparse-attention architecture is maturing into a production-ready product. MiniMax-M3 has not yet published comparative benchmarks against direct rivals, but its spec sheet positions it as a serious contender for anyone who needs a million tokens of context without paying a closed-model premium.