Inference Brew

UkisAI Launches Swift-1.5-Qwen3.8-27b Model

00:00 / --:--

← Back to home

UkisAI Launches Swift-1.5-Qwen3.8-27b Model

1. UkisAI Launches Swift-1.5-Qwen3.8-27b Model

Building on the On-Policy Distillation techniques announced on September 14th, UkisAI has officially released Swift-1.5-Qwen3.8-27b. This variant is fine-tuned to maximize token efficiency for fast, low-thinking tasks. Comparative testing shows the model's IQ4_XS quantization outperforms Unsloth's Q4_K_S in low-thinking scenarios, while maintaining high performance on local hardware like the NVIDIA 3090.

  • • UkisAI has released Swift-1.5-Qwen3.8-27b, a model variant optimized for token efficiency.
  • • This release follows the team's September 14th announcement regarding On-Policy Distillation for Qwen 3.8 27B.
  • • The model demonstrates superior performance in low-thinking tasks compared to Unsloth's Q4_K_S quantization.
  • • Community testing confirms local execution on an NVIDIA 3090 at 67 tokens per second.

Provides developers with a production-ready, token-efficient model variant that builds on UkisAI's previously announced optimization research.

SOURCES

2. Enterprise Coding Agents Update Pricing, Indemnity, and Rebranding

The enterprise AI coding assistant landscape has seen major shifts in ownership, pricing, and legal protections. Cognition has acquired the Windsurf IDE, rebranding it as Devin Desktop. On the pricing front, GitHub Copilot has transitioned to a usage-based credit model, while also relaxing its IP indemnity requirements for unmodified code. Meanwhile, AWS Kiro is offering uncapped copyright indemnity, contrasting with Cursor's MSA which excludes indemnity from its liability cap.

  • • Cognition has acquired the Windsurf IDE and rebranded it as Devin Desktop as of June 2, 2026.
  • • GitHub Copilot transitioned to a usage-based AI credit model on June 1, 2026, priced at $0.01 per credit.
  • • Microsoft's GitHub Copilot IP indemnity policy no longer requires additional mitigations for unmodified code.
  • • AWS Kiro offers uncapped copyright indemnity, while Cursor's Master Service Agreement excludes indemnity from its fee cap.
  • • GitHub Enterprise Cloud now supports EU or US data residency for a 10% increase in AI credit consumption.
  • • Cursor's Privacy Mode prevents data retention and training, though requests still process through Cursor's backend.

Helps developers and enterprise teams navigate the shifting legal, privacy, and pricing landscapes of major AI coding assistants.

SOURCES

3. Google Research Introduces Massive Sound Embedding Benchmark (MSEB)

Google Research has launched the Massive Sound Embedding Benchmark (MSEB), a framework designed to standardize the evaluation of audio encoders. Rather than relying on a single headline metric, MSEB scores models across four distinct tasks: classification, clustering, retrieval, and segmentation. Developers can integrate models like Whisper, wav2vec, CLAP, EnCodec, and SoundStream by implementing a simple three-method contract, with the framework handling all batching, validation, and statistical calculations.

  • • Google Research has introduced the Massive Sound Embedding Benchmark (MSEB) for evaluating sound encoders.
  • • The framework tests models across classification, clustering, retrieval, and segmentation tasks.
  • • Encoders must implement a simple three-method contract: _setup, _check_input_types, and _encode.
  • • MSEB supports integration with popular audio models including Whisper, wav2vec, CLAP, EnCodec, and SoundStream.
  • • The benchmark handles batching, validation, and statistics, and can run on a free CPU runtime without dataset downloads.

Provides a standardized framework to evaluate and score sound encoders across multiple tasks, helping developers choose the best audio embedding model for their application.

SOURCES

4. Public MCP Server Released for Canadian Privacy Law Data

Privacy researcher Mohammad Movahedi has released a free, public Model Context Protocol (MCP) server and REST API containing Canadian privacy law data. Operating over Streamable HTTP transport, the server exposes five specialized tools that allow AI agents to search enforcement actions from Quebec's access-to-information commission, look up terms in a 263-word privacy glossary, and evaluate compliance against an 11-point Quebec Law 25 readiness checklist.

  • • A free, public Model Context Protocol (MCP) server has been released providing access to Canadian privacy law data.
  • • The server utilizes Streamable HTTP transport and is hosted publicly at https://movahedi.ca/mcp.
  • • It includes five tools for searching enforcement actions, looking up glossary terms, and checking Quebec Law 25 readiness.
  • • An accompanying REST API is available with quotas of 2,000 daily calls for anonymous users and 10,000 for free API key holders.

Gives developers an out-of-the-box MCP server to query Canadian privacy laws and Quebec Law 25 compliance requirements directly within their AI agents.

SOURCES

5. Llama.cpp Accelerates Prompt Lookup Drafting by 42x

The llama.cpp repository has received a major performance update, accelerating its Prompt Lookup Drafting implementation by 42x. This optimization dramatically speeds up speculative decoding, allowing local models to draft and verify tokens much faster during inference without requiring additional hardware.

  • • Prompt Lookup Drafting in llama.cpp has been updated to deliver a 42-fold increase in speed.
  • • The optimization significantly improves speculative decoding performance for local model runs.

Drastically reduces latency for local LLM inference by accelerating speculative decoding via prompt lookup drafting.

SOURCES

6. Logit-Bias Penalties on Overthinking Words Boost Qwen Accuracy

An intriguing experiment demonstrates that applying logit-bias penalties to 49 specific "overthinking" markers—such as "perhaps," "maybe," "wait," and "actually"—can significantly improve Qwen model accuracy. Testing on the Qwen3.5-4B-GGUF model using the MATH-500 dataset showed that the BF16 format's accuracy jumped from 74% to 84% while simultaneously reducing reasoning tokens by 19.4%. Even the highly quantized Q2_K format saw its accuracy double, suggesting that discouraging circular reasoning steps can yield both better answers and lower token costs.

  • • Applying logit-bias penalties to 49 "overthinking" words (like "perhaps", "maybe", "wait") improved Qwen model accuracy.
  • • In tests on Qwen3.5-4B-GGUF, the BF16 format's accuracy rose from 74% to 84% on the MATH-500 dataset.
  • • The technique reduced reasoning tokens by 19.4% for the BF16 model and 11.5% for the Q2_K model.
  • • Accuracy for the highly quantized Q2_K format doubled from 12% to 24%.
  • • The penalties discourage unnecessary reasoning steps, though results are currently based on a single preliminary test.

Allows developers to boost reasoning model accuracy and reduce token costs by applying simple logit-bias penalties to overthinking words during inference.

SOURCES

7. SSD Streaming Engine Runs 177B MoE Models on Consumer GPUs

A breakthrough inference engine for Mixture-of-Experts (MoE) models allows developers to run massive models locally by streaming experts directly from an SSD when they exceed VRAM and RAM limits. In tests, the engine achieved 9-10 tokens per second decoding the 176.9B parameter Qwen3.8-Flash-Next NVFP4 model using just a single 16GB RTX 5060 Ti GPU, 32GB of RAM, and a Gen5 NVMe SSD. The system utilizes GCLOCK eviction, lookahead prefetch, and native NVFP4 matrix multiplications running directly on Blackwell tensor cores, though it is currently limited to RTX 50-series GPUs, Windows 11/WSL2, and greedy decoding.

  • • Developers have built an inference engine that streams MoE experts from an SSD to run models that exceed VRAM and RAM limits.
  • • The engine achieved 9-10 tokens per second decoding a 176.9B Qwen3.8-Flash-Next NVFP4 model on a single 16GB RTX 5060 Ti.
  • • The system uses 20 GiB of RAM and 99 GiB of SSD storage, leveraging GCLOCK eviction and lookahead prefetch.
  • • It features native NVFP4 matrix multiplications running directly on Blackwell tensor cores without unpacking.
  • • Currently, the engine is limited to RTX 50-series (Blackwell) GPUs, Windows 11 or WSL2, and greedy decoding.
  • • The project includes an OpenAI-compatible server with built-in chat and tool-calling support.

Enables developers to run massive 177B parameter models locally on budget consumer GPUs by streaming model weights directly from high-speed SSDs.

SOURCES

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.

Inference Brew in your inbox

5 minutes a day. Free, unsubscribe anytime.