Inference Brew

SpaceXAI Upgrades Frontier Model to Grok 4.6

00:00 / --:--

← Back to home

SpaceXAI Upgrades Frontier Model to Grok 4.6

1. SpaceXAI Upgrades Frontier Model to Grok 4.6

SpaceXAI has released Grok 4.6, the successor to the Grok 4.5 model. Designed for long-running agents, coding, and knowledge work, the new model features improved safeguards and training on regenerated SFT trajectories. It is now available via the Grok Build platform, Cursor, OpenRouter, Vercel, and Cloudflare.

  • • Grok 4.6 succeeds Grok 4.5, scoring 61 on the Artificial Analysis Intelligence Index.
  • • Pricing remains consistent with the previous version at $2 per million input tokens and $6 per million output tokens.
  • • Features a 500,000-token context window, function calling, and structured outputs.
  • • Available immediately via Grok Build, Cursor, OpenRouter, Vercel, and Cloudflare.

Developers gain access to an updated, agent-optimized frontier model that matches GPT-5.6 Sol performance, building on the capabilities established in the previous Grok 4.5 release.

2. Alibaba Releases Open-Weights Qwen3.8-2.4T-A95B-FP8 MoE Model

Alibaba has fulfilled its promise to release open weights for the Qwen3.8-Max model, now available as Qwen3.8-2.4T-A95B-FP8. This massive mixture-of-experts model features 95 billion active parameters, a native 262K context window, and a mandatory thinking mode for reasoning. It is now available for self-hosting via major inference frameworks.

  • • Qwen3.8-2.4T-A95B-FP8 features 2.4 trillion total parameters with 95 billion activated parameters across 512 experts.
  • • The model supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.
  • • It requires a mandatory thinking mode where reasoning is generated and enclosed in tags before output.
  • • Users can adjust reasoning depth using the reasoning_effort parameter (xhigh, medium, low).
  • • The model is compatible with Transformers, vLLM, and SGLang inference frameworks.

Developers can now self-host the frontier-class Qwen3.8-Max model, which was previously restricted to API access on QwenCloud.

3. New Open-Weight Models from DeepSeek, NVIDIA, and Mistral Join the Ecosystem

The open-weight model landscape continues to expand rapidly. Building on the recent release of Meta's Muse Glimmer, this week sees the addition of DeepSeek-V4-Flash-0731, NVIDIA's Nemotron-3.5-Lightning-30B-A3B and VoiceChat-11B, Liquid AI's LFM2.5-2.6B, and Mistral's Shieldstral-1.0-3B, providing developers with a diverse set of new tools for local deployment.

  • • DeepSeek-V4-Flash-0731 is a 304B MoE model under the MIT license.
  • • NVIDIA released Nemotron-3.5-Lightning-30B-A3B and the full-duplex VoiceChat-11B.
  • • Liquid AI released LFM2.5-2.6B, supporting 131k context.
  • • Mistral released Shieldstral-1.0-3B, a multimodal guardrail model.

Developers now have an even wider selection of specialized, permissive-license models for local deployment, low-latency inference, and agentic workflows beyond the recently released Muse Glimmer.

SOURCES

4. Upstage Launches Solar Pro 4 Reasoning Model

Upstage has launched Solar Pro 4, a proprietary flagship reasoning model that demonstrates significant performance improvements in agentic and long-context tasks. The model features a massive context window and a high maximum output limit, making it suitable for complex, multi-step reasoning workflows.

  • • Upstage released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, up from Solar Pro 3's score of 14.
  • • Pricing is set at $0.30 per million input tokens and $1.20 per million output tokens via Upstage's first-party API.
  • • The model features a 512K token context window and a 128K token maximum output limit.
  • • It is available via Upstage's first-party API, OpenRouter, and TimelyRouter.
  • • Solar Pro 4 reduced its hallucination rate to 24% and increased its Terminal-Bench v2.1 score to 57%.

Developers can leverage a low-cost reasoning model with a massive context window and improved agentic performance for high-volume tasks.

SOURCES

5. Liquid AI Releases LFM2.5-VL-3B On-Device Vision Model

Liquid AI has released LFM2.5-VL-3B, a multimodal variant of its LFM2.5 hybrid model family. Combining an LFM2.5-2.6B language backbone with a SigLIP2 NaFlex 400M vision encoder, the model is optimized for local, on-device deployment, offering high-throughput inference within a small memory footprint.

  • • Liquid AI released LFM2.5-VL-3B, a 3.1B parameter vision-language model with a 32K context window.
  • • The model runs locally, achieving 228 tokens/s on an Apple M5 Max and 20 tokens/s on a Galaxy S26 Ultra within 3GB of memory.
  • • It features a SigLIP2 NaFlex 400M vision encoder and supports 16 languages.
  • • The model is optimized for near-realtime object detection, OCR with layout annotation, and on-device translation.
  • • It is available on Hugging Face and can be tested using the atomic.chat mobile application.

Developers can deploy a fast, lightweight vision model locally on consumer hardware for real-time object detection, OCR, and translation.

SOURCES

6. GitHub Upgrades Copilot with MAI-Code-1.1-Flash

Building on the initial launch of the MAI-Code-1-Flash model in June 2026, GitHub has now released MAI-Code-1.1-Flash. This updated model offers 25% greater token efficiency and is priced at one-quarter of the original model's cost, while providing improved performance on coding tasks.

  • • MAI-Code-1.1-Flash succeeds the original MAI-Code-1-Flash model released in June 2026.
  • • The new model offers 25% greater token efficiency and costs 75% less than its predecessor.
  • • Performance improvements include a 22% gain on Terminal-Bench 2.1 and a 15% improvement on .NET tasks.
  • • The model is now available for use within GitHub Copilot.

Developers using GitHub Copilot will benefit from faster, higher-quality code generation and improved CLI performance at a significantly lower cost than the previous iteration.

SOURCES

7. Cohere Releases North-Micro-Vision-Instruct Model

CohereLabs has released North-Micro-Vision-Instruct, an open-weight vision-language model designed for research and development tasks. The model features a 2B-parameter language backbone and a 400M-parameter vision encoder, supporting native-resolution image processing and multi-image inputs.

  • • North Micro Vision Instruct is a 2.4B-parameter model combining a 2B language backbone and a 400M SigLIP 2 vision encoder.
  • • The model is released under the Apache 2.0 license and is available on Hugging Face.
  • • It supports native-resolution image processing, multi-image inputs, and a 128K-token language context window.
  • • The model is designed for research, prototyping, fine-tuning, and document understanding.
  • • It does not support tool calling or agentic workflows, and has limited math and code-generation capabilities.

Developers can self-host a lightweight, permissive-license vision model for prototyping, document understanding, and visual QA.

SOURCES

8. Mistral Launches Regional Endpoints and Priority Tier

Mistral is advancing AI sovereignty by providing enterprises with greater control over AI models and infrastructure. The company has made its Regional Endpoints generally available, allowing customers to select Europe or the US for inference processing, and has introduced a Priority Tier offering committed service levels and custom rate limits.

  • • Mistral Regional Endpoints are generally available, allowing customers to select Europe or the US for inference processing.
  • • The Mistral Priority Tier is in public preview, offering committed service levels, custom rate limits, and an uptime SLA.
  • • Mistral is expanding its platform to support third-party open models, starting with Z.ai's GLM-5.2.
  • • The company introduced European Compute Units (ECUs) to convert multi-year commitments into infrastructure access.

Developers can now guarantee in-region inference processing in Europe or the US and secure committed service levels for mission-critical workloads.

SOURCES

9. LiteLLM Supply-Chain Attack Exposes Credentials for Over 2,500 Organizations

A major supply-chain attack on the open-source tool LiteLLM has exposed credentials for over 2,500 organizations, including Microsoft, Amazon, Cisco, Samsung, and Salesforce. The breach occurred because victims downloaded compromised versions of LiteLLM from the Python Package Index (PyPI), originating from a prior supply-chain attack on the Trivy vulnerability scanner.

  • • A supply-chain attack on LiteLLM exposed credentials, including cloud keys, repository tokens, and AI provider keys, for over 2,500 organizations.
  • • The breach occurred because users downloaded compromised versions of LiteLLM from the Python Package Index (PyPI).
  • • The compromise originated from a prior supply-chain attack on the Trivy vulnerability scanner, which also affected KICS and the Telnyx Python SDK.
  • • Stolen credentials were leaked during a 40-minute window in March, and the legitimacy of the data has been confirmed by security researchers.

Developers using LiteLLM must immediately audit their environments, rotate exposed API keys, and ensure they are not running compromised versions of the package.

SOURCES

10. Cursor Rebrands 'Origin' to 'Cursor Review' for Automated PR Pipeline

Following the initial launch of the 'Origin' Git-compatible forge in June, Cursor is now transitioning the platform out of closed beta under the new name 'Cursor Review'. The updated platform will feature a 'Codebase' tab for repository management and a 'Review' tab to facilitate automated pull request workflows between humans and AI agents.

  • • Cursor is rebranding its 'Origin' platform to 'Cursor Review'.
  • • The platform is transitioning out of a closed partner beta.
  • • New features include a 'Codebase' tab for repository management and a 'Review' tab for automated PR pipelines.
  • • The public rollout is expected as early as this week.

This transition marks the move of Cursor's agent-native Git forge from a closed beta into a broader, feature-focused platform for collaborative code review.

SOURCES

11. Speculative Decoding Speeds Up Muse Glimmer 30B on Apple Silicon

A developer has implemented speculative decoding for Meta's Muse Glimmer 30B model in the mlx-dspark project on Apple Silicon. The implementation achieves significant speedups across math, code, and chat tasks while ensuring the output remains byte-identical to normal decoding, preserving model quality.

  • • Speculative decoding in the mlx-dspark project increased Muse Glimmer 30B speed on an M4 Pro from 8.2 to up to 26 tokens/second.
  • • Speedup factors for the 8-bit model were 3.27x for math, 2.5x for code, and 2.22x for chat.
  • • The output remains byte-identical to normal decoding, ensuring no loss in model quality.
  • • The 8-bit run peaks at 40GB of memory usage, while the 4-bit build provides a 1.7x speedup requiring 18GB of memory.
  • • The project is open-source and available on GitHub.

Developers running local models on Apple Silicon can dramatically reduce latency without any loss in output quality.

SOURCES

12. New Study Finds Token Reduction Tools Can Increase LLM Costs

Following earlier reports that token-compression tools like rtk and headroom offer only marginal real-world savings due to the 'denominator effect,' a new study by PointFive suggests the impact may be worse than previously understood. Evaluating 2,908 Claude Code sessions across 103 tasks, the study found that these tools—often marketed as cost-saving measures—can actually increase total LLM costs by up to 46.4% depending on the model and task configuration.

  • • PointFive evaluated 2,908 Claude Code sessions and 103 tasks to measure the real-world cost impact of token reduction tools.
  • • The study found that these tools can increase LLM costs by up to 46.4%, contradicting marketing claims of up to 90% savings.
  • • This finding expands on previous June 2026 benchmarks that identified minimal 3.7% savings and operational risks associated with these tools.

Developers should move beyond skepticism of advertised savings and actively benchmark their token reduction strategies, as these tools may now be identified as a source of increased API expenditure rather than a cost-optimization solution.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.