Moonshot releases massive MoE weights but they need 2 TB RAM to run

Moonshot, the Chinese AI firm, released the weights for Kimi K3 this week, a mixture-of-experts (MoE) model with 2.8 trillion parameters and 104 billion active units per forward pass. On paper it is the largest model currently available for download. In practice, running it locally requires a system with 2 TB of RAM, putting it beyond the reach of most developers and small research teams. The weights are open, but the training data and code have not been published, so it is “open weights,” not open source.
OpenAI cut prices in version 5.6: Luna fell fivefold to $0.2 per million input tokens and $1.2 per million output tokens (about 0.74 and 4.4 shekel respectively), Terra was reduced by 20 percent, and Sol added a Fast mode that runs 2.5× faster for double the cost. The move comes as competition over inference cost sharpens, and it is likely intended to protect market share against cheaper alternatives.
DeepSeek V4 Flash, API version 0731, outperforms GLM 5.2 in benchmarks, supports Codex and remains cheaper than Luna both per token and per task compute. A Pro version is expected in early August. At the same time, Astra, OpenAI’s new multi-agent model, solved 10 open math problems using Lean and priced under $2,000 per million tokens on Sol. Automatic proof generation in mathematics remains expensive, but this figure is far lower than anything seen before.
ByteDance’s Seedance 2.5 now creates clips up to 180 seconds, adds Blender integration and point-level editing, but is priced at $6 for 30 seconds, making it the most expensive video-generation service on the market. The capabilities are impressive, yet the cost per video second restricts it to budgeted productions rather than daily experimentation.