aiexpert
Home / News / Brief
Breaking · Jul 24, 2026, 01:35 AM · 3 sources

Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing

Together AI now hosts DeepSeek V4 Pro, the 1.6-trillion-parameter mixture-of-experts model, with 512K-token context and three controllable reasoning modes (Non-Think, Think High, Think Max). The model uses hybrid Compressed Sparse Attention and Heavily Compressed Attention, reducing single-token inference FLOPs to 27% and KV cache to 10% versus DeepSeek V3.2 at million-token context. Pricing is $2.10 per 1M input tokens, $0.20 per 1M cached input tokens, and $4.40 per 1M output tokens.

DeepSeek V4 Pro benchmarks at 93.5% on LiveCodeBench, 90.1% on GPQA Diamond, 80.6% on SWE-Bench Verified, and 83.5% on MRCR 1M comprehension. The 1.6T parameter count activates only 49B parameters per forward pass, giving the model frontier-level knowledge capacity while keeping inference costs comparable to a 49B dense model. Cached input pricing provides 90% cost reduction for repeated analysis over the same large context—critical for code agents, document intelligence, and long-horizon agentic workflows.

For teams building long-context reasoning systems, DeepSeek V4 Pro on Together AI removes the operational burden of running a trillion-parameter MoE locally while maintaining serverless flexibility or moving to dedicated, reserved capacity for production SLA guarantees. Workloads like repository analysis, policy comparison, and multi-step agentic decision-making can now leverage million-token context at hosted-inference pricing. Architects evaluating open-weight inference hosting should benchmark V4 Pro's reasoning modes and cached-input cost profile against Groq, Fireworks, and direct cloud provider inference endpoints.

Sources

Everything this brief rests on
  1. 01 Primary source together.ai
  2. 02 DeepSeek V4 Pro model card, architecture details huggingface.co
  3. 03 Together AI pricing ($2.10/$0.20/$4.40 per 1M tokens) aipricing.guru