Inference Brew

Moonshot AI Announces Kimi K3 2.8-Trillion-Parameter MoE Model

00:00 / --:--

← Back to home

Moonshot AI Announces Kimi K3 2.8-Trillion-Parameter MoE Model

1. Moonshot AI Announces Kimi K3 2.8-Trillion-Parameter MoE Model

Moonshot AI has announced Kimi K3, a 2.8-trillion-parameter sparse Mixture-of-Experts model. Built on Kimi Delta Attention and Attention Residuals architectures, the model activates 16 of 896 experts to improve decoding speed and scaling efficiency. Kimi K3 is currently accessible via the Kimi API, with full model weights scheduled for release on July 27, 2026.

  • Moonshot AI announced Kimi K3, a 2.8-trillion-parameter sparse Mixture-of-Experts model with weights scheduled for release on July 27, 2026.
  • The model features native visual understanding, an always-on 'thinking mode' reasoning capability, and a 1-million-token context window.
  • It utilizes Kimi Delta Attention and Attention Residuals to improve decoding speed and scaling efficiency.
  • API pricing is set at $3.00 per million input tokens ($0.30 for cache hits) and $15.00 per million output tokens.
  • In a 48-hour demonstration, Kimi K3 autonomously designed a functional 4-square-millimeter chip using open-source electronic design automation tools.

Developers can build on a massive, frontier-class open-weights model with native multimodal support and long-context reasoning.

2. GPT-5.6 Sol Leads Web Design Arena Benchmark

Building on the global preview of the GPT-5.6 model family, OpenAI's flagship Sol model has now achieved first place on Design Arena's Web Design Arena. This performance milestone highlights the model's capability to recognize and compress AI design anti-patterns, marking a significant improvement over the GPT-5.5 model, which currently ranks 18 positions lower.

  • GPT-5.6 Sol has reached the number one ranking on the Web Design Arena benchmark.
  • The model demonstrates a significant performance lead over GPT-5.5, which currently ranks 18 positions lower.
  • The benchmark results validate the model's ability to effectively identify and compress common AI design anti-patterns.

Developers testing the GPT-5.6 Sol model in the current global preview can now leverage its validated top-tier performance for generating personalized, high-quality web designs.

SOURCES

3. OpenLLM-France Releases Luciole-23B-Instruct-1.1

LINAGORA and the OpenLLM-France consortium have released Luciole-23B-Instruct-1.1, a fine-tuned and aligned version of the open-source, multilingual Luciole-23B-Base model. Training was conducted on the Jean Zay supercomputer in three phases, including supervised fine-tuning with and without thinking traces, followed by Direct Preference Optimization (DPO). The collection is licensed under Apache 2.0 and includes 8B and 1B parameter versions.

  • Luciole-23B-Instruct-1.1 is a fine-tuned and aligned version of the open-source, multilingual Luciole-23B-Base model.
  • The model was developed by LINAGORA and the OpenLLM-France consortium, with funding from BPI France.
  • Training was conducted on the Jean Zay supercomputer using supervised fine-tuning with and without thinking traces, followed by DPO alignment.
  • The collection is released under an Apache 2.0 license and includes 23B, 8B, and 1B parameter versions.

Developers have access to a new Apache 2.0 licensed multilingual model fine-tuned on math, science, coding, and RAG tasks.

SOURCES

4. OpenAI Details Root Cause and Mitigation for GPT-5.6 File Deletion Bug

Building on earlier reports of file deletion issues in GPT-5.6, OpenAI has identified the root cause as an environment variable error where the model mistakenly overrides the $HOME directory while attempting to define a temporary path. To mitigate the risk, OpenAI is updating developer messaging and providing guidance on safer permission modes while preparing a full post-mortem.

  • The file deletion is caused by the model incorrectly overriding the $HOME environment variable.
  • OpenAI is mitigating the issue by updating developer messages and recommending safer permission modes.
  • A detailed post-mortem on the incident is expected in the coming days.

Understanding the specific technical failure allows developers to better configure their environments and apply necessary sandboxing to prevent data loss until a permanent fix is deployed.

SOURCES

5. LM Studio Launches Bionic AI Agent for Open Models

LM Studio has launched Bionic, a new AI agent application designed for open models. The platform supports both local model execution and cloud-based execution via the LM Studio Secure Cloud, committing to Zero Data Retention. For coding tasks, Bionic can inspect local codebases, debug, and perform inline diffs, with support for models such as GLM 5.2 and Kimi K2.7 Code.

  • LM Studio launched Bionic, a separate AI agent application supporting local model execution and cloud execution via LM Studio Secure Cloud.
  • Bionic features zero data retention and does not train on user data.
  • For coding, it can inspect local codebases, debug, and perform inline diffs using models like GLM 5.2 and Kimi K2.7 Code.
  • It includes a voice keyboard utilizing Mistral AI's Voxtral model for local, real-time transcription.
  • The platform provides a sandboxed environment for document tasks with automatic checkpoints for version control.

Developers can use a dedicated agent environment to run open models locally or via a secure cloud for coding, debugging, and document tasks with zero data retention.

SOURCES

6. Modal Rebuilds Sandbox Platform to Support 1M Concurrent Environments

Modal has rebuilt its sandbox platform to support millions of concurrent sandboxes and tens of thousands of sandbox creations per second. The new architecture removes central bottlenecks by replacing global coordination with a horizontally scalable, asynchronous design using Redis streams. Median sandbox start times are now less than half a second, requiring only two network hops and one CPU operation.

  • Modal rebuilt its sandbox platform to support millions of concurrent sandboxes and tens of thousands of creations per second.
  • The new architecture replaces global coordination with a horizontally scalable, asynchronous design using Redis streams.
  • Median sandbox start times are now less than half a second, requiring only two network hops and one CPU operation.
  • Modal benchmarked the platform by successfully creating 1 million sandboxes in under one minute.
  • The new platform is currently available in Beta for users to opt into.

Developers running agent workloads can scale to massive numbers of concurrent, isolated execution environments with sub-second startup times.

SOURCES

7. Perplexity Launches SPACE Ephemeral Sandbox Platform for Agents

Perplexity AI has launched SPACE, a sandbox platform designed to secure AI agents performing sensitive tasks. The platform utilizes ephemeral sandboxes that are automatically destroyed once a task is completed. SPACE manages and protects credential access through a Control Plane and Node-level Services, offering security features like credential isolation, rolling snapshots, and encrypted storage.

  • Perplexity AI launched SPACE, a sandbox platform designed for secure and efficient AI agent execution.
  • The platform uses ephemeral sandboxes that are automatically destroyed upon task completion.
  • It isolates and protects credentials using a Control Plane and Node-level Services.
  • Security features include credential isolation, rolling snapshots, and encrypted storage.
  • SPACE supports both on-premises and offline operations.

Developers can run sensitive agentic tasks in secure, isolated, and automatically destroyed environments to prevent credential exposure and data leaks.

SOURCES

8. 1Password Integrates with Claude for Secure Credential Autofill

1Password has launched a browser integration that allows the Anthropic Claude chatbot to use stored security credentials to complete tasks. The integration utilizes a zero-exposure security framework that injects credentials through a secure channel, preventing the AI model from viewing actual passwords or MFA codes. Users must approve each request via a biometric prompt, and 1Password automatically locks down access to only the specific credentials granted.

  • 1Password launched a browser integration allowing the Anthropic Claude chatbot to use stored credentials to complete tasks.
  • A zero-exposure security framework injects credentials through a secure channel, preventing the AI model from viewing passwords or MFA codes.
  • Users must approve each credential request via a biometric prompt, and 1Password locks down access to only the specific credentials granted.
  • The feature is available for 1Password users on Mac across business, family, and individual plans, requiring both desktop apps and browser extensions.
  • Future updates plan to add support for payment cards and identity details.

Developers can safely grant Claude agents access to secure credentials without exposing raw passwords or MFA codes to the underlying LLM.

SOURCES

9. ReasonGate Open-Sources Explainable Prompt Injection Guard

ReasonGate has been released as an explainable, model-agnostic security gate designed to protect LLM applications by inspecting user prompts, retrieved context, and model outputs. Written in pure Python with zero dependencies, the core tool wraps any prompt-to-string function and includes detectors for normalization, injection, indirect injection, multi-turn risk, and output leakage.

  • ReasonGate is an explainable, model-agnostic security gate licensed under Apache-2.0.
  • The core tool is written in pure Python with zero dependencies and wraps any prompt-to-string function.
  • It includes detectors for normalization, injection, indirect injection, multi-turn risk, and output leakage.
  • A policy engine fuses signals to decide whether to allow, flag, or block requests, generating structured, machine-readable audit records.
  • An optional enterprise add-on provides embedding-based ML detection, while the core remains rule-only.

Developers can wrap any LLM or RAG pipeline in a lightweight, self-contained security layer to inspect prompts and generate structured audit records.

SOURCES

10. Granola Launches MCP Integration for Meeting Notes

Granola has launched a Model Context Protocol (MCP) integration that exposes meeting notes directly to AI tools like Claude and ChatGPT. This integration allows developers to automate workflows, such as updating CRM systems or organizing tasks in Linear, asynchronously. By using MCP, developers can process meeting data without requiring AI bots to join live calls.

  • Granola launched an MCP integration that exposes meeting notes to AI tools like Claude and ChatGPT.
  • The integration enables automated workflows, such as updating CRM systems or organizing tasks in Linear.
  • It processes notes asynchronously, eliminating the need for AI bots to join live calls.
  • Granola is offering the first month of service free with the code TLDR1MO.

Developers can connect meeting notes directly to their AI workflows, enabling automated tasks like updating CRMs or Linear without inviting bots to live calls.

SOURCES

11. Libretto PR Agents Automatically Fix Failing Playwright Scripts

Saffron Health has released Libretto PR agents, a free, open-source TypeScript library designed to maintain Playwright browser automations. The tool allows an agent to automatically open GitHub pull requests to fix Playwright scripts when they fail. It connects to a failed browser session via CDP and uses an execution tool to inject Playwright and JavaScript into the page to inspect the failure.

  • Saffron Health released Libretto PR agents, a free, open-source TypeScript library for maintaining Playwright browser automations.
  • The tool automatically opens GitHub pull requests to fix Playwright scripts when they fail.
  • It integrates into existing scripts with a single line of code and connects to failed browser sessions via CDP.
  • The library supports user-provided LLM API keys and works with any browser provider, including self-hosted options.

Developers can integrate this library with a single line of code to let AI agents autonomously debug and repair broken end-to-end tests.

SOURCES

12. Atlassian Integrates AI Agents Directly into Jira

Atlassian has introduced new capabilities in Jira that allow users to assign tasks directly to AI agents, including Claude, Cursor, or GitHub Copilot. This integration passes relevant task context directly from the Jira interface to the agent. Atlassian claims that these new capabilities provide 44% better agent output.

  • Atlassian added capabilities to Jira for assigning tasks directly to AI agents including Claude, Cursor, or GitHub Copilot.
  • The integration passes relevant task context directly from the Jira interface to the agent.
  • Atlassian claims the new capabilities provide 44% better agent output.

Developers can receive and execute Jira tasks directly within their AI-native coding environments with full context.

SOURCES

13. Developer Automates Backlog Management with Claude Code

A developer has automated their backlog management process by setting up a self-improving loop using Claude Code. The automated system is capable of triaging tasks, decomposing large tasks into smaller ones, implementing the code, running tests, and opening pull requests. The entire pipeline was run for a cost of approximately $110 per month.

  • A developer automated their backlog management process using Claude Code.
  • The system triages tasks, decomposes large tasks into smaller ones, implements code, runs tests, and opens pull requests.
  • The entire self-improving pipeline was run for a cost of $110 per month.

Developers can implement similar automated loops to triage tasks, write code, run tests, and open pull requests autonomously.

SOURCES

14. Patter SDK Enables Local Prototyping of Voice Agents

A new tutorial demonstrates how to use the Patter SDK to build a voice-agent workflow for a restaurant booking use case. The SDK allows developers to prototype and test voice agents in a self-contained environment, defining dynamic caller variables, registering callable tools, and applying output guardrails. The tutorial also provides a template for deploying the tested logic to live telephony using Twilio and OpenAI Realtime.

  • The Patter SDK tutorial outlines building a voice-agent workflow for a restaurant booking use case.
  • The workflow supports dynamic caller variables, callable tools, output guardrails, and simulated speech-to-text/text-to-speech.
  • The system tracks latency and cost metrics and includes a deterministic evaluation harness for regression testing.
  • Prototyped logic can be deployed to live telephony using Twilio and OpenAI Realtime.

Developers can prototype and test voice agents in a self-contained environment before deploying them to live telephony infrastructure like Twilio and OpenAI Realtime.

SOURCES

15. ReactBench v1 Released to Evaluate Coding Agents

ReactBench v1 has been released as an evaluation framework designed to test the performance of coding agents. The framework specifically evaluates how well coding agents handle realistic React development tasks, providing developers with a standardized benchmark for frontend agent capabilities.

  • ReactBench v1 is an evaluation framework designed specifically for coding agents.
  • The framework tests agent performance on realistic React development tasks.

Developers can use this framework to benchmark and compare the performance of different coding agents on frontend React tasks.

SOURCES

16. Open Interpreter Runs Coding Agents Locally

Open Interpreter provides an open-source environment for running coding agents locally. The software is designed to test both web and native application interfaces, allowing developers to automate and evaluate UI interactions directly on their local machines.

  • Open Interpreter runs coding agents locally on a user's machine.
  • The software is designed to test both web and native application interfaces.

Developers can execute local coding agents to automate and test user interfaces without relying on cloud-hosted environments.

SOURCES

17. Benchmarks Reveal 6x Speedup for DFlash Speculative Decoding in llama.cpp

Building on the recent integration of DFlash speculative decoding support in llama.cpp, new benchmarks using the Qwen 3.6 27B model on an NVIDIA RTX PRO 6000 GPU show significant performance gains. Combining DFlash with n-gram lookup drafters achieved a 6.01x speedup for iterative coding tasks, with performance reaching 7.5x during code maintenance. These results quantify the practical benefits of the previously announced DFlash support, noting that n-gram lookup tables stored in host RAM incur no additional VRAM cost.

  • Benchmarks combining DFlash and n-gram lookup drafters in llama.cpp achieved a 6.01x speedup (321.5 tok/s) on an 18-turn iterative coding task.
  • Speedups reached 7.5x (385 tok/s) during maintenance turns involving edits to existing code.
  • N-gram drafters store lookup tables in host RAM, incurring zero additional VRAM cost.
  • The system remains output-lossless at greedy temperature because the target model verifies every drafted token.
  • DFlash is recommended for structured coding tasks, while MTP is recommended for chat and creative writing.

These benchmarks validate the performance potential of the DFlash speculative decoding integration, showing that developers can achieve substantial speedups for coding tasks without increasing VRAM usage.

SOURCES

18. DeepSeek V4 Flash Performance Gains 300% Following llama.cpp Updates

Building on the previously reported integration of DeepSeek V4 support and quantized KV cache capabilities in llama.cpp, recent updates between versions b9986 and b10034 have significantly improved inference performance. Users running the 98GB DeepSeek-V4-Flash-UD-Q2_K_XL model on consumer hardware—specifically an AMD Ryzen 5 9600X with an NVIDIA RTX 4060 Ti—are now seeing generation speeds increase from 2 to 7 tokens per second. These gains were achieved using a context size of 131,072 and layer-based split mode.

  • Generation speeds for the 98GB DeepSeek-V4-Flash-UD-Q2_K_XL model improved from 2 to 7 tokens per second.
  • Performance gains are attributed to llama.cpp updates between versions b9986 and b10034.
  • Testing was conducted on consumer hardware: AMD Ryzen 5 9600X, 138GB RAM, and an NVIDIA RTX 4060 Ti (16GB VRAM).
  • The configuration utilized a 131,072 context window and layer-based split mode.

These performance improvements make running massive, high-parameter models locally on budget-friendly consumer hardware significantly more practical for developers.

SOURCES

19. EU Orders Google to Open Android and Search to Rival AI Assistants

The European Commission has issued legally binding measures under the Digital Markets Act requiring Google to open Android and Google Search to competing AI platforms. Under the ruling, Google must allow rival AI assistants to access system features, device hardware, and app interactions on Android by July 2027. Additionally, Google must share search data with competing search engines and AI chatbots starting in January 2027.

  • The EU ordered Google to share search data by January 2027 and implement Android interoperability changes by July 2027.
  • The Android ruling mandates that Google allow rival AI assistants to access system features, device hardware, and app interactions.
  • The Search ruling requires Google to share search data with competing search engines and AI chatbots.
  • Non-compliance could result in European Commission fines of up to 10 percent of Google's annual worldwide turnover.

Developers of alternative AI assistants like ChatGPT or Claude will gain deep system-level integration on Android devices in Europe, ending Gemini's exclusive access.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.