Inference Brew

Anthropic Releases Claude Opus 5.5 with Lower Pricing and Enhanced Safety

00:00 / --:--

← Back to home

Anthropic Releases Claude Opus 5.5 with Lower Pricing and Enhanced Safety

1. Anthropic Releases Claude Opus 5.5 with Lower Pricing and Enhanced Safety

Following the release of Claude Opus 5 and the Fable 5.1 series, Anthropic has launched Claude Opus 5.5. The new flagship model features a 20% reduction in base token pricing to $4 per million input and $20 per million output, alongside a 60% cut in cache read costs. Opus 5.5 is 30% faster than its predecessor and introduces stricter safety safeguards, including an 85% reduction in boundary circumvention attempts and automated routing for sensitive cybersecurity and biology queries.

  • • Claude Opus 5.5 is the latest addition to the Claude 5 family, featuring a 1M token context window.
  • • Base pricing is reduced by 20% to $4 per million input and $20 per million output tokens, with cache read costs cut by 60% to $0.20 per million.
  • • The model is over 30% faster than Opus 5 and currently holds the top spot on the Artificial Analysis Intelligence Index with a score of 58.
  • • Enhanced safety protocols reduce boundary circumvention attempts by 85% compared to previous versions.
  • • Sensitive cybersecurity and biology requests are automatically re-routed to older model versions (Opus 4.8 and Opus 5, respectively).

Developers can now access a more cost-effective and secure flagship model that currently leads the Artificial Analysis Intelligence Index, building on the capabilities established in the earlier Claude 5 releases.

2. OpenAI Expands GPT-6 Lineup with Sol and Luna Models

OpenAI has expanded its GPT-6 model family with the general availability of GPT-6 Sol and GPT-6 Luna, following the initial release of the agentic GPT-6 Astra model earlier this month. These new additions provide developers with mainstream and budget-friendly alternatives to the high-end Astra model, featuring significant price reductions compared to the GPT-5.6 series and improved hallucination rates.

  • • OpenAI expanded the GPT-6 series with the release of Sol and Luna models.
  • • GPT-6 Sol is priced at $2/M input and $10/M output tokens.
  • • GPT-6 Luna is priced at $0.10/M input and $0.50/M output tokens.
  • • Both models feature improved hallucination rates and maintain the 90% prompt cache discount introduced with the GPT-6 architecture.
  • • GPT-6 Sol outperforms Claude Opus 5 on the AutomationBench benchmark at a lower cost.

Developers now have a broader range of GPT-6 models to choose from, allowing them to balance performance, cost, and reasoning capabilities across the new architecture.

3. Xiaomi Expands MiMo-V2.6 Series with Flash Omnimodal Model

Expanding on the recent launch of MiMo-V2.6-Pro, Xiaomi has released the Flash omnimodal model, designed for coordinating agents in 3D scene construction. The company has also open-sourced the technical report, training environments, and reinforcement learning code for the V2.6 series. Both models feature a 1M token context window and support text, image, speech, and video inputs.

  • • Xiaomi released the Flash omnimodal model alongside the previously announced MiMo-V2.6-Pro.
  • • The models are designed to coordinate agents for constructing and testing interactive 3D scenes.
  • • Xiaomi open-sourced the technical report, training environments, and reinforcement learning code for the V2.6 series.
  • • Both models support text, image, speech, and video inputs with a 1M token context window.

Developers now have access to the full MiMo-V2.6 suite, including the new Flash model and comprehensive open-source training resources for building interactive 3D agentic workflows.

4. StepFun Launches Step 5 Preview with 1M Context Window

StepFun has launched Step 5 Preview, a new flagship model featuring 600 billion total parameters (27 billion active) and a 1M token context window. The model accepts text, image, and video inputs, and matches the performance of Kimi K3 (max) on the Artificial Analysis Intelligence Index while offering a 2.8x lower cost per task. It is currently available via StepFun's first-party API at $1 per million input tokens and $2.70 per million output tokens, with an open weights release planned for October 15th.

  • • StepFun released Step 5 Preview, a flagship model with 600B total parameters and 27B active parameters.
  • • The model supports a 1M token context window and accepts text, image, and video inputs.
  • • Step 5 Preview matches the performance of Kimi K3 (max) on the Artificial Analysis Intelligence Index while costing 2.8 times less per task.
  • • Pricing is $1 per million input tokens and $2.70 per million output tokens, with cached input at $0.05 per million tokens.
  • • The model is available via StepFun's API, with an open weights release scheduled for October 15th.

Developers get a highly cost-effective flagship model for long-context multimodal tasks, with an open-weights release scheduled for next month.

5. AntLing Open-Sources Ming-Image-0.1-Design 6B Model Family

AntLing has open-sourced the Ming-Image-0.1-Design model family, consisting of the 6B Ming-Image-0.1-Design and 6B Ming-Image-0.1-Design-Layer models. Designed specifically for UI/UX tasks, the base model currently ranks first among open-weight models on the Artificial Analysis UI/UX Design leaderboard. The release also packages two open-source agent skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill.

  • • AntLing open-sourced the 6B parameter Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer models.
  • • The release includes two open-source agent skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill.
  • • Ming-Image-0.1-Design ranks first among open-weight models on the Artificial Analysis UI/UX Design leaderboard.

Developers can self-host a top-ranked open-weight model specifically tuned for generating UI designs and editable presentations.

SOURCES

6. StepFun Releases StepAudio 3 ASR with 1.7% Word Error Rate

StepFun has released StepAudio 3 ASR, a new non-streaming Speech to Text model available through its API. The model achieves a 1.7% Word Error Rate (WER) on the AA-WER Index, tying for first place with Alibaba's Fun-Realtime-ASR-preview and marking a significant improvement over StepAudio 2.5's 4.7% WER. However, at 88x real-time speed, it trails faster alternatives like MAI-Transcribe-2 (374x) and is priced at a premium of $0.40 per hour.

  • • StepFun released StepAudio 3 ASR, a non-streaming Speech to Text model available via its API.
  • • The model achieved a 1.7% Word Error Rate (WER) on the AA-WER Index, tying with Alibaba's Fun-Realtime-ASR-preview.
  • • StepAudio 3 ASR transcribes at a speed factor of 88x real time, trailing competitors like MAI-Transcribe-2 (374x) and Grok Voice Transcribe 2.0 (154x).
  • • The model is priced at $0.40 per hour ($6.67 per 1,000 minutes), making it the most expensive of the five most accurate models.

Developers building voice features get a highly accurate speech-to-text option, though they must weigh its high cost against alternative fast transcribers.

SOURCES

7. Qwen-Image 2.1 Uses 9B Prompt Rewriter for Image Generation

Qwen-Image 2.1 utilizes a 9B parameter prompt rewriter to expand user requests before generating images. This rewriter requires 20 GB of memory in bf16 format, uses a 1,700-word system prompt, and processes approximately 1,600 tokens of "thinking" before outputting. For resource-constrained environments, a compressed 0.8B parameter version of the model has been released to enable execution on laptops.

  • • Qwen-Image 2.1 utilizes a 9B parameter prompt rewriter to expand user requests before generating images.
  • • The 9B prompt rewriter requires 20 GB of memory in bf16 format and uses a 1,700-word system prompt.
  • • The rewriter thinks for approximately 1,600 tokens before generating an answer.
  • • A compressed 0.8B parameter version of the model is available to run on laptops.

Understanding the prompt expansion architecture of Qwen-Image 2.1 helps developers optimize their image generation pipelines and hardware requirements.

SOURCES

8. AWS Integrates Strands Framework into New Agent Harness

Following the open-source release of the Strands Agents framework in June, AWS has now launched the Strands Agent Harness. This new offering transitions the framework into a managed, ready-to-run environment that includes built-in web searching, command execution, file editing, and task delegation capabilities, allowing developers to deploy agents by simply supplying an LLM.

  • • AWS has launched the Strands Agent Harness, an evolution of the open-source Strands Agents framework.
  • • The harness provides a ready-to-run environment for web searching, command execution, and file editing.
  • • It includes built-in memory and task delegation to helper agents.
  • • Developers only need to provide the underlying LLM to begin execution.

Developers can now leverage the established Strands framework within a fully managed AWS environment, simplifying the deployment of agentic workflows with built-in tool execution and delegation.

SOURCES

9. DigitalOcean Launches Managed Agents in Public Preview

DigitalOcean has launched Managed Agents in public preview, providing a serverless runtime environment designed specifically for running autonomous agents like Claude Code, Codex, or custom LangGraph workflows. The service supports over 75 open and proprietary models, consolidates billing, and provides a single governed endpoint for agent tools. To optimize infrastructure costs, the managed runtime automatically pauses when agents are idle.

  • • DigitalOcean Managed Agents has entered public preview.
  • • The service runs Claude Code, Codex, or custom LangGraph agents in a runtime environment that pauses when idle.
  • • It supports over 75 open and proprietary models.
  • • The platform provides a single governed endpoint for agent tools and consolidates billing.

Developers can deploy and run autonomous agents in a secure, managed environment that automatically pauses when idle to save costs.

SOURCES

10. Rabbit Launches OS3 Agentic Operating System with BYOK Model

Rabbit has announced OS3, an agentic operating system designed to run across Windows, Mac, and Linux devices without requiring the company's R1 hardware (which has ceased manufacturing). OS3 operates on a "Bring Your Own Key" (BYOK) model, allowing users to integrate their own API keys from providers like OpenAI or Anthropic. The system runs a local agent node on the user's machine to execute software tasks, write code, and debug, automatically coordinating the necessary files, apps, and models.

  • • Rabbit announced OS3, an agentic operating system that runs across Windows, Mac, and Linux, as well as the R1 gadget.
  • • OS3 uses a "Bring Your Own Key" (BYOK) model, allowing users to supply their own API keys from OpenAI or Anthropic.
  • • The platform uses a local agent node on the host machine to execute software tasks, write code, and debug.
  • • Rabbit has ceased manufacturing the R1 hardware and plans to release a "Cyberdeck" device for "vibe coding" in the coming months.

Developers and power users can run local agent nodes to execute software tasks and write code using their own API keys.

11. Firecrawl Raises $75M and Launches Alexandria Cloud Service

Web scraping startup Firecrawl has raised $75 million in a Series B funding round led by Smash Ventures. Alongside the funding, Firecrawl has launched Alexandria, a new cloud service that integrates web scraping data with third-party and curated datasets, including scientific abstracts and code files. Alexandria exposes this data through a single API, eliminating the need for developers to build custom connectors, while leveraging Firecrawl's agent-optimized scraping features like smart wait and automated form submission.

  • • Firecrawl raised $75 million in a Series B funding round led by Smash Ventures.
  • • The platform provides web scraping tools optimized for AI agents, featuring smart wait and automated user actions (scrolling, clicking).
  • • Firecrawl launched Alexandria, a new cloud service that integrates web data with third-party and curated datasets (scientific abstracts, code files).
  • • Alexandria provides data through a single API, eliminating the need for custom connectors.

Developers building AI agents get a more robust, funded scraping platform and a unified API to access web data alongside curated code and scientific datasets.

SOURCES

12. Jev Catalog Curates Tools and Integrations for Decision Models

Jev is a specialized tool designed to answer data queries by selecting options, providing yes/no probabilities, or scoring items on a scale rather than generating free-form text. To integrate Jev, developers use official JavaScript, TypeScript, or Python SDKs configured with a typesafe-sdk API key. A curated catalog of 36 reviewed projects—including tools for browser automation and coding agent support—is updated every four hours to showcase community integrations.

  • • Jev is a tool that answers questions by selecting options, providing yes/no probabilities, or scoring items instead of generating text.
  • • The project requires the typesafe-sdk and a TYPESAFE_API_KEY.
  • • Official support is provided through JavaScript, TypeScript, and Python SDKs.
  • • The curated catalog features 36 picks, including tools for browser automation, coding agent support, and gaming.

Developers can explore a curated catalog of 36 projects and integrations built around structured decision models.

SOURCES

13. Optimizing Rust Code Performance Using Agentic LLMs

A developer has documented a highly effective method for using agentic LLMs to iteratively optimize Rust code, achieving performance speedups of 2x to 20x. The workflow relies on a structured "Ur-Prompt" and a set of rules defined in an AGENTS.md file to guide LLMs through benchmarking with the criterion crate while preventing cheating behaviors. The process employs GPT-5.6 Luna subagents to explore optimization hypotheses and includes a refactoring step to reduce source lines of code by at least 20% without performance regressions.

  • • The author developed an "Ur-Prompt" and rules in AGENTS.md to guide LLMs in benchmarking and optimizing Rust code.
  • • The process achieved performance speedups ranging from 2x to 20x across various domains.
  • • Benchmarking is performed using the criterion crate, with constraints forbidding unsafe code and parallelized runs.
  • • The workflow uses GPT-5.6 Luna subagents to explore optimization hypotheses and includes a refactoring step to reduce source lines of code by 20%.
  • • Recent frontier models like GPT-6 Astra demonstrated the ability to find algorithmic reimplementations for an additional 2x-3x speedup.

Developers can adopt structured prompting rules and subagent workflows to automate the benchmarking and optimization of their local codebases.

SOURCES

14. Artificial Analysis Expands Speech Arena with Multilingual Support and Pronunciation Benchmark

Artificial Analysis has expanded its Controlled Voice Arena to include nine new languages (Japanese, Mandarin, Hindi, Spanish, German, French, Portuguese, Vietnamese, and Arabic) and introduced a Pronunciation Robustness benchmark. This update builds on the platform's existing Speech Arena, which previously tracked model performance in English and other primary languages. The new benchmark evaluates TTS models on 454 challenging sentences, with Google Gemini 3.1 Flash TTS currently leading at 88.1% accuracy. Cartesia's Sonic 3.6, previously noted as a top performer on the platform, now leads in seven of the nine newly added languages.

  • • Artificial Analysis added nine languages to its Controlled Voice Arena, including Japanese, Mandarin, and Arabic.
  • • A new Pronunciation Robustness benchmark evaluates TTS accuracy across 454 sentences.
  • • Google Gemini 3.1 Flash TTS leads the new pronunciation benchmark with 88.1% accuracy.
  • • Cartesia's Sonic 3.6 ranks first in seven of the nine newly added languages.
  • • Kokoro 82M v1.0 is the fastest model tested at 242 characters per second but scored 55.5% on pronunciation robustness.

Developers can now evaluate TTS models for global applications using expanded language support and specific pronunciation accuracy metrics, extending the utility of the existing Artificial Analysis benchmarking suite.

15. JevBench Released to Standardize Evaluation of Jev-Class Models

As the ecosystem for Jev-class models—specialized systems designed for structured decision-making—continues to grow, developers have released JevBench. This open-source benchmarking tool provides a reproducible harness of 534 English decisions, allowing for the standardized evaluation of intelligence, calibration, speed, and cost across different Jev-style implementations.

  • • JevBench provides a standardized, open-source harness for evaluating Jev-class models.
  • • The tool measures accuracy, calibration, latency, and cost across 534 English decision tasks.
  • • This release supports the growing ecosystem of Jev-style models, including commercial, open-source, and DIY implementations.
  • • The harness, artifacts, and scoring code are available on GitHub.

Standardized benchmarking allows developers to compare the performance of various Jev-class models and DIY implementations, facilitating better selection and optimization for structured decision-making tasks.

SOURCES

16. TypeSafe AI Releases Jev2 and SemIf Models for If-Statement Optimization

Following the September 16 launch of the Jev framework, TypeSafe AI has introduced Jev2 and SemIf, two new models specifically engineered to streamline if-then decision-making. These models build on the original Jev architecture, demonstrating a 99% reduction in costs and an accuracy increase from 47% to over 80% in testing, further refining the economics of AI-driven conditional logic.

  • • Jev2 and SemIf are new models designed to optimize if-then decision-making.
  • • These models follow the initial Jev framework launch from September 16.
  • • Testing shows a 99% cost reduction and accuracy improvement from 47% to over 80%.
  • • The models aim to further transform the economics of AI production systems.

These updates provide developers with more specialized, efficient tools to replace complex conditional logic, significantly lowering API costs compared to the original Jev release.

SOURCES

17. Drop Open-Source Rootless Linux Sandbox Adds gVisor Support

Drop is an open-source, rootless Linux sandbox designed to run third-party programs in isolated environments, making it highly useful for developers executing untrusted LLM-generated code. The tool leverages Linux namespaces to enforce isolation, providing a writable, disposable home directory while mounting configuration files as read-only. The latest update introduces optional gVisor user-space kernel support, offering enhanced protection against host kernel exploits without requiring root privileges.

  • • Drop is a rootless Linux sandbox designed to run third-party programs in isolated environments.
  • • The tool uses Linux namespaces (user, mount, network, PID, IPC, cgroup) to enforce isolation.
  • • It provides a writable, disposable home directory and mounts configuration files as read-only.
  • • The latest update adds optional gVisor user-space kernel support to protect against host kernel exploits.

Developers building agent runtimes or executing untrusted LLM-generated code can use Drop to secure their host systems without requiring root privileges.

SOURCES

18. Cisco Talos Releases CAIRN Framework to Analyze AI-Integrated Malware

Cisco Talos has released an open-source framework called the Cognitive Artifact Intelligence Research Network (CAIRN) to classify and analyze AI-integrated malware. Using CAIRN, researchers identified CLOSEDQUORUM, a Windows credential- and crypto-stealing malware that autonomously determines its actions by polling four LLMs (DeepSeek, Qwen, Mistral, and Google Gemini). The release highlights a growing trend of attackers operationalizing LLM APIs, such as the previously documented LAMEHUG malware which communicated with Qwen2.5-Coder via Hugging Face.

  • • Cisco Talos released the open-source Cognitive Artifact Intelligence Research Network (CAIRN) framework to classify and analyze AI-integrated malware.
  • • Researchers used CAIRN to identify CLOSEDQUORUM, a Windows malware that autonomously polls DeepSeek, Qwen, Mistral, and Gemini to determine its actions.
  • • CLOSEDQUORUM is designed to steal login credentials and cryptocurrency.
  • • In July 2025, a phishing campaign was found using LAMEHUG malware to communicate with Qwen2.5-Coder-32B-Instruct via a Hugging Face API.

Developers and security teams can use CAIRN to analyze emerging AI-driven threats, such as malware that uses LLM APIs to coordinate attacks.

SOURCES

19. Aikido Altar Open-Weight Model Released for On-Premise Pentesting

Aikido Altar has been released as an open-weight AI model designed to enhance the pentesting capabilities of the Aikido Machine. Built specifically for sovereign security, the model runs entirely within a customer's own local infrastructure, allowing developers to perform automated security testing without exposing sensitive code or network data to external cloud APIs.

  • • Aikido Altar is an open-weight AI model designed to improve the pentesting capabilities of the Aikido Machine.
  • • The model is built to operate entirely within a customer's own infrastructure.

Security-conscious developers can deploy a specialized pentesting model locally to scan their infrastructure without sending sensitive data to external APIs.

SOURCES

20. Qwen Image 2.1 Fast FP8 Supported in Unsloth Studio

Users have reported successful local execution of the Qwen Image 2.1 (Fast FP8) model under 10GB of VRAM using Unsloth Studio. To access this high-performance, low-memory image generation option, developers are advised to update Unsloth Studio to the latest version.

  • • Qwen Image 2.1 (Fast FP8) can be run under 10GB of VRAM on Unsloth Studio.
  • • The model generates premium quality images with high performance.
  • • Users need to update Unsloth Studio to access the latest model options.

Developers can run high-quality image generation locally on consumer GPUs with low memory requirements.

SOURCES

21. GGUF Quantizations Now Available for K2-Horizon Models

Building on the September 3 launch of the K2-Horizon open-weight model family, quantized GGUF versions are now available on Hugging Face. These versions, ranging from 0.9B to the 36B MoVA model, allow for local execution, though they currently require a specific fork of llama.cpp while official support is pending.

  • • Quantized GGUF versions of the K2-Horizon model series are now available on Hugging Face.
  • • Supported models include K2-Horizon-MoVA-36B, 32B, 7B, 3.7B, and 0.9B.
  • • Local execution requires a specific fork of llama.cpp as official support is still in progress.

This update enables developers to run the K2-Horizon models locally on consumer hardware, expanding on the initial release of the model weights.

SOURCES

22. Dynamic Quantiser Tool Optimizes GGUF Files

Dynamic Quantiser is a new open-source tool designed to create custom, high-quality GGUF quantizations of local models. By minimizing whole-model cosine deviation using Dinkelbach's algorithm and a separable Lagrangian inner loop, the tool determines the optimal quantization level for each tensor. While it is less performant than Unsloth dynamic 3.0 on pure text tasks, it is reported to be more accurate than standard K-quants.

  • • Dynamic Quantiser is an open-source GitHub repository for creating custom GGUF quantizations.
  • • The tool minimizes whole-model cosine deviation to improve model accuracy using Dinkelbach's algorithm.
  • • It offers three algorithms: dynamic quantization, cosine deviation minimization, and disk size minimization.
  • • The tool is more accurate than standard K-quants, though less performant than Unsloth dynamic 3.0 on pure text tasks.
  • • Initial table building takes about 25 minutes for a 27B model on a Ryzen 5 3600 processor.

Developers running local models can generate highly accurate custom quantizations that outperform standard K-quants.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.