1. OpenAI Opens GPT-Live-1 to API Developers
OpenAI has expanded access to its GPT-Live-1 model, moving it from a ChatGPT-exclusive feature to a developer-facing API product. Priced at $0.05 per minute, the model enables developers to integrate full-duplex speech capabilities—allowing for simultaneous listening and speaking—into their own applications. This release follows the model's initial consumer rollout in July, providing developers with native interruption handling and real-time reasoning capabilities for voice agents.
- • GPT-Live-1 is now available in the OpenAI API at $0.05 per minute.
- • The model supports full-duplex speech, allowing simultaneous listening and speaking.
- • Includes 12 voice options and native interruption handling.
- • Enables background reasoning or actions during live voice interactions.
- • Early testing indicates an 80% reduction in interruptions compared to previous turn-based systems.
Developers can now build highly natural, real-time voice interfaces that handle interruptions gracefully and perform complex reasoning mid-conversation using the same technology previously limited to ChatGPT.
2. OpenAI Launches GPT Image 2.5 Flare and Sunburst Models
OpenAI has updated its image generation suite with the release of GPT Image 2.5 Flare and Sunburst. Flare is built for low-latency, everyday image generation, while Sunburst is tailored for detailed creative work and precise editing control. Both models are priced at $30 per million image output tokens—matching the pricing of GPT Image 2—and are immediately accessible via the OpenAI API. Sunburst has taken the top spot on the Artificial Analysis Image Editing leaderboard, demonstrating significant improvements in handling complex layouts, text rendering, and structural edits.
- • OpenAI released GPT Image 2.5 Flare and Sunburst, five months after GPT Image 2.
- • Flare is optimized for low-latency everyday generation, while Sunburst is designed for high-quality creative work and tight editing control.
- • Both models support Text to Image and Image Editing, and are priced at $30 per million image output tokens.
- • Flare and Sunburst hold the top two positions on the Artificial Analysis image leaderboards.
- • Sunburst leads in 8 of 10 Image Editing use cases, showing major gains in complex composition, diagrams, and creator content.
- • The models are available through the OpenAI API and integrated into ChatGPT and Codex.
Developers can integrate faster, higher-quality image generation and precise composition editing into their applications without paying a premium over previous-generation models.
3. Cohere Releases North Small Translate 218B MoE Model
Cohere has released North Small Translate, a 218-billion-parameter sparse Mixture-of-Experts (MoE) model optimized for high-quality machine translation across 50 languages. Activating 25 billion parameters per token, the model features a decoder-only architecture with 128 experts and supports a generous 16K input/output context window. Developers can access the model via Cohere's API or self-host the 4-bit quantized version on a single B200 or dual H100 setup, making it a viable option for local, high-throughput translation pipelines.
- • North Small Translate is a sparse Mixture-of-Experts (MoE) model with 218B total and 25B active parameters.
- • The model supports translation across 50 languages with a 16K input and 16K output context window.
- • It is available via Cohere's API, for non-commercial self-hosting, or under a commercial license.
- • The 4-bit quantized version can be run locally on a single B200 or two H100 GPUs.
- • An agentic variant uses a multi-pass workflow to correct errors, scoring 84.36 on Cohere's WMT26 evaluation.
Developers can deploy a highly accurate, open-weights translation model locally on standard enterprise hardware or access it via API for multilingual applications.
4. Orukeet 25-Language Speech Recognition Model Released
Orukeet has been released as a high-performance, 25-language automatic speech recognition (ASR) model built on top of the NVIDIA Parakeet TDT 0.6B v3 architecture. By replacing half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels, Orukeet achieves a 10.6% relative reduction in word error rate (WER) across 25 FLEURS languages compared to the original Parakeet model. This open-weights model offers developers a highly accurate, multi-accent speech-to-text solution for local deployment.
- • Orukeet is a 25-language speech recognizer built from the NVIDIA Parakeet TDT 0.6B v3 model.
- • The model replaces half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels.
- • Orukeet achieves a 1.46% word error rate (WER) on LibriSpeech test-clean, compared to 1.53% for Parakeet.
- • Across 25 FLEURS languages, Orukeet achieves a pooled WER of 9.85%, representing a 10.6% relative reduction over Parakeet.
- • The model is trained on multilingual and multi-accent data, using LibriSpeech test-other for final adaptation.
Developers building voice and speech features can adopt a highly accurate, open-weights speech-to-text model that reduces word error rates across multiple languages.
5. Qwen3.8-27B-Humanlike-Chat Model Released
To combat the overly verbose and polished tone characteristic of standard AI assistants, a developer has released Qwen3.8-27B-Humanlike-Chat. Fine-tuned using a rank-256 LoRA on the abliterated Qwen3.8-27B base model, the training leveraged a dataset of over 125,000 obfuscated human-to-human messages. The resulting model naturally generates shorter, less formal, and more realistic conversational responses without requiring complex system prompting. It is currently available as open weights on Hugging Face and via a rate-limited, OpenAI-compatible API endpoint.
- • The model was trained to address the overly helpful, polished, and verbose tone typical of standard AI assistants.
- • It was trained using a dataset of 125,217 obfuscated human-to-human messages across 1,396 conversations.
- • The training involved a rank-256 LoRA applied to the huihui-ai/Huihui-Qwen3.8-27B-abliterated base model.
- • The model produces shorter, less polished responses without requiring specific system prompts.
- • It is available via Hugging Face, a demo space, and a rate-limited OpenAI-compatible API endpoint.
Developers building conversational agents can bypass overly verbose, robotic assistant templates and deploy a model that naturally mimics realistic human-to-human dialogue.
6. Sakana AI Expands Fugu Orchestration Platform with Fugu Max and Fugu Ultra v2
Building on the Fugu orchestration platform previously expanded with the Fugu-Cyber endpoint, Sakana AI has launched two new models: Fugu Max and Fugu Ultra v2. These models are available via a single, OpenAI-compatible API, allowing developers to route tasks across a pool of models for improved cost-efficiency or complex reasoning. Fugu Max is optimized for lower costs, while Fugu Ultra v2 focuses on high-capability multi-step reasoning. Note that these models are not available for self-hosting and remain excluded from the EU/EEA.
- • Fugu Max and Fugu Ultra v2 are new additions to the Fugu orchestration platform.
- • Fugu Max is priced at $2/M input and $6/M output tokens, targeting cost-efficiency.
- • Fugu Ultra v2 is designed for complex reasoning, scoring 74.3 on the DeepSWE benchmark.
- • The models are available via a hosted, OpenAI-compatible API.
- • Like previous Fugu endpoints, these are not available for self-hosting and are excluded from the EU/EEA.
This update provides developers with additional specialized routing options within the Fugu ecosystem, enabling more granular control over cost and reasoning performance.
7. Octen Search Debuts with Ultra-Low Latency and Cost
Octen Search has entered the search API market with a strong focus on speed and cost-efficiency for AI agents. Debuting third on the Artificial Analysis Search Index, the service clocks an average query latency of just 0.2 seconds—five times faster than Perplexity Search variants—and completes complex search tasks in 16.9 seconds. Priced at a highly competitive $1 per 1,000 queries, Octen Search provides developers with a high-performance, low-cost option for grounding LLMs and agents in real-time web data.
- • Octen Search debuted third on the Artificial Analysis Search Index with a score of 77.
- • The service achieved the fastest Time per Task at 16.9 seconds, beating the model-only baseline of 22.6 seconds.
- • It averages 0.2 seconds per query, which is 5x faster than Perplexity Search variants.
- • Octen Search is priced at $1 per 1,000 search queries ($0.058 per Search Index task).
- • The benchmark was conducted using the Stirrup open-source agent harness with a fixed base model.
Developers building search-enabled agents can drastically reduce latency and API costs by switching to a search provider optimized for speed.
8. Devin Fusion Multi-Model Coding Agent Benchmarked and Launched
Building on the recent integration of the SWE-2 model, Cognition has formally launched Devin Fusion. This multi-model agent architecture pairs a frontier model—such as Claude Fable 5.1 or GPT-6 Astra—with a SWE-2 sidekick model to balance capability and execution costs. The system has now debuted on the Artificial Analysis Coding Agent Index, providing developers with benchmarked data on speed and cost-efficiency for agentic workflows.
- • Devin Fusion is the first multi-model coding agent featured on the Artificial Analysis Coding Agent Index.
- • The configuration pairing Claude Fable 5.1 (xhigh) with SWE-2 (medium) achieved a score of 62 on the Coding Agent Index v1.5.
- • The GPT-6 Astra (xhigh) and SWE-2 (medium) configuration scored 59, while being 43% cheaper and 31% faster than the Claude Fable setup.
- • Cognition published a launch blog post detailing the technical breakdown of the Fusion model.
Developers can now evaluate the performance and cost trade-offs of Devin Fusion's multi-model configurations using standardized benchmarks from the Artificial Analysis Coding Agent Index.
9. Anthropic Introduces Plugin Evals for Claude Code
Anthropic has added a native plugin evaluation workflow to Claude Code, starting in version 2.1.269. This new feature allows developers to measure the performance of custom plugins by comparing them against a baseline where the plugin is disabled, generating a clear delta score. The framework supports six grader types—four of which are free to run, while two utilize paid LLM calls to act as judges. With the new 'claude-code eval' command, developers can automatically generate test suites and integrate them directly into CI pipelines to prevent regressions in agent skills.
- • Anthropic has introduced a new plugin evaluation workflow for Claude Code (v2.1.269 or later).
- • The workflow compares plugin performance against a no-plugin baseline to calculate a delta score.
- • The system supports six grader types, including four free graders and two paid LLM-judge graders.
- • The 'claude-code eval' command can automatically generate initial test suites and graders for a plugin.
- • Evaluations can be integrated directly into CI pipelines to act as a quality gate for new skills.
Developers building custom plugins or skills for Claude Code can now systematically test, benchmark, and gate their code using automated, multi-grader evaluations.
10. Litelm: A Lightweight, Low-Dependency Alternative to LiteLLM
Developers looking to trim down their application dependencies can now use litelm, a lightweight alternative to the popular LiteLLM library. Written in just 2,900 lines of code with only two dependencies ('openai' and 'httpx'), litelm focuses strictly on core routing, message translation, streaming, tool use, and embeddings while omitting heavy features like proxy servers and caching. The API is designed to mirror LiteLLM directly, allowing for drop-in replacement, and includes full async support alongside verified integration with DSPy.
- • litelm is a minimal LLM routing library consisting of ~2,900 lines of code and only two dependencies (openai and httpx).
- • The API mirrors LiteLLM, allowing developers to switch by simply updating their import statements.
- • It supports routing to 19 providers using a standard provider/model-name syntax and includes async variants like acompletion.
- • The library maps provider errors to a custom exception hierarchy (e.g., ContextWindowExceededError).
- • Currently in alpha, the library has verified support for DSPy integration across seven execution paths.
Developers can replace bloated routing libraries with a highly optimized, minimal dependency wrapper that maintains API compatibility and native async support.
11. CodeFinetuner Simplifies Local Autocomplete Model Fine-Tuning
CodeFinetuner has launched as an end-to-end pipeline designed to help developers fine-tune small, local code autocomplete models (such as Qwen2.5-Coder-3B) on their own private codebases. Installable via 'uv tool install codefinetuner', the tool supports training on both Mac (MPS) and NVIDIA (CUDA) hardware, utilizing Unsloth to minimize VRAM usage. The pipeline automates the entire process—from tree-sitter parsing and LoRA fine-tuning to evaluation and GGUF conversion—producing models that can be dropped directly into local editor extensions like llama.vim or llama.vscode.
- • CodeFinetuner is an end-to-end pipeline for LoRA fine-tuning of small code autocomplete models on specific codebases.
- • The tool supports training on both Mac (MPS) and NVIDIA (CUDA) hardware, with optional Unsloth integration.
- • The workflow handles tree-sitter parsing, LoRA fine-tuning, evaluation (CodeBLEU/perplexity), and GGUF conversion.
- • Resulting GGUF models are compatible with local editor tools like llama.vim and llama.vscode.
- • The tool can be installed instantly via 'uv tool install codefinetuner'.
Developers can easily customize local autocomplete models to their private codebases, running them efficiently on Mac or NVIDIA hardware.
12. Google Launches Cloud Developer Plugin Based on Agent Plugins 1.0.0 Standard
Following the release of the Agent Plugins 1.0.0 open standard in August, Google has launched the Google Cloud Developer Plugin. This new toolset provides installable bundles that equip AI coding agents with specialized skills for interacting with, configuring, and managing Google Cloud infrastructure, marking a practical application of the recently established plugin specification.
- • Google released the Google Cloud Developer Plugin, implementing the Agent Plugins 1.0.0 standard.
- • The plugin provides installable bundles of tools for managing Google Cloud resources.
- • It enables AI agents to perform deployment, management, and troubleshooting tasks within Google Cloud environments.
This release demonstrates the adoption of the Agent Plugins 1.0.0 standard, allowing developers to integrate standardized, portable cloud-management capabilities directly into their AI agent workflows.
13. Google Research Releases ToolGrad Framework for Tool-Use Data Generation
Researchers from Google, the University of Tokyo, RIKEN, and Tohoku University have released ToolGrad, an open-source framework designed to generate high-quality training data for tool-use and function-calling models. By inverting the traditional pipeline—building a verified tool-use chain first and then generating the corresponding user query—ToolGrad achieves a 99.8% data generation pass rate on ToolBench. Using this method, the researchers fine-tuned a Gemma-3 12B model that scored 83.1 on the Berkeley Function Calling Leaderboard, actually outperforming its teacher model, Gemini 2.5 Flash-Lite. The framework is available as a PyPI package under an Apache-2.0 license.
- • ToolGrad inverts standard data generation by constructing a verified tool-use chain first, then generating the matching user query.
- • The framework improved the data generation pass rate on ToolBench from 63.8% to 99.8% compared to depth-first search.
- • The researchers used ToolGrad to fine-tune Gemma-3 models (1B, 4B, 12B) using Gemini 2.5 Flash-Lite.
- • The resulting ToolGrad-12B model scored 83.1 on the Berkeley Function Calling Leaderboard, outperforming its teacher model.
- • ToolGrad code is released under an Apache-2.0 license, with models, datasets, and a PyPI package available.
Developers can generate high-quality function-calling datasets and fine-tune small, highly efficient local models that outperform frontier APIs at tool use.
14. Llama-Manager Enables Dynamic Mid-Generation Model Reconfiguration
A new tool called llama-manager offers developers a way to dynamically reconfigure local models running on llama.cpp without interrupting active token generation. By preserving the KV cache during reconfigurations, llama-manager eliminates the need to re-process prompts when toggling speculative decoding, moving components to the CPU, or adjusting KV cache quantization. In beta testing on a single 32GB GPU, the tool successfully expanded the context window of a Qwen3.8-27B model from 167,680 tokens to 262,144 tokens using dynamic memory strategies.
- • llama-manager is a wrapper for a llama.cpp fork that enables dynamic model configuration after loading.
- • The tool supports hot reloading of models mid-token generation, preserving the KV cache to avoid re-processing prompts.
- • It allows users to dynamically toggle speculative decoding, move components to the CPU, and update KV cache quantization.
- • When running Qwen3.8-27B-UD-Q4_K_XL on a 32GB GPU, it expanded the context window from 167,680 to 262,144 tokens.
- • The project is currently in beta and is restricted to single-GPU inference.
Developers running local models can dynamically adjust quantization, speculative decoding, and offloading strategies on the fly, maximizing context length on limited GPU memory.
15. Fine-Tuning Qwen 3 4B on 100 Logic Puzzles Boosts Math Benchmark by 31%
A striking demonstration of targeted fine-tuning shows that training the Qwen 3 4B Base model on a tiny dataset of just 100 zebra puzzles yields a 31% performance boost on the MATH-500 benchmark. The entire training run was completed in 6.5 minutes on a single H100 or H200 GPU. This recipe, complete with a shared reproduction notebook, highlights how developers can achieve massive reasoning gains in small, local models using highly curated, domain-specific datasets rather than massive compute budgets.
- • Fine-tuning Qwen 3 4B Base on 100 zebra puzzles resulted in a 31% improvement on the MATH-500 benchmark.
- • The training process took only 6.5 minutes on a single H100 or H200 GPU.
- • A complete reproduction notebook for the fine-tuning process has been made publicly available.
Developers can apply highly targeted, ultra-small datasets to dramatically improve the reasoning capabilities of small, local models with minimal compute costs.