New AI models as they ship, from frontier LLMs and open-weight releases to image, video and audio models, with technical analysis and independent benchmarks.
LegalOn says it cut estimated daily Codex costs by 65% while maintaining development speed. It attributed the savings to matching Astra, Sol, and Luna models to different tasks and managing budgets strategically.
OpenAI safety researchers Tomek Korbak, Mikita Balesni and Jasmine Wang say they were fired after raising concerns about AI safety and the company’s work with evaluator METR; OpenAI reportedly says they mishandled confidential information. The roundup also reports releases including OpenAI’s GPT-6.1 Sol Ultrafast and Anthropic’s Claude Haiku 5.5, alongside new research and benchmark findings on agent safety, model performance and infrastructure.
Hugging Face author Yvrj Sharma says its ML-intern agent helped build six models and LoRAs over several days, with projects costing about $103 in total compute. Examples include a 0.8B image-prompt rewriter that runs on a CPU and achieved 99.7% valid outputs, a citrus-disease vision model that raised accuracy from 14.9% to 52.8% on 335 test images, and a four-step image generator that improved GenEval from 0.509 to 0.536. The agent plans and runs training, evaluation and publication jobs after getting a budget approved.
Mistral has released a preview of Mistral Large 4, a 1-trillion-parameter model with 49 billion active parameters, trained on 3,800 NVIDIA Grace Blackwell GPUs. The API preview offers “none” and “high” reasoning modes; Mistral says open weights will arrive at the end of the month, and the model scores 38 on Artificial Analysis, up from 9 for Mistral Large 3.
Anthropic released Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with rates five times higher above that; blogger Simon Willison notes its tokenizer uses about 25% more tokens than Haiku 4.5. Anthropic also halved Sonnet 5.5 cache-read prices and announced monthly API credits of up to $100 for Max 5x, $200 for Max 20x, and $500 pooled for Team subscribers; credits expire monthly.
OpenAI released 722 manuscripts from an unreleased internal math model, covering 372 groups of results drawn from an evaluation of about 4,000 problems. The company says the work used roughly three hours of ChatGPT Pro compute per result; commentators highlight a claimed quasi-Riemann hypothesis result, but the manuscripts and proofs have not been independently verified.
The page appears to be a short post by Simon Willison titled “llm-mistral 0.16,” tagged with LLM, Mistral and reasoning. The provided text contains no details about what changed in version 0.16.
Simon Willison argues that embedding models should be available under open licenses because replacing a discontinued proprietary model can force users to recalculate millions of stored vectors. He welcomes EmbeddingGemma 2’s Apache 2.0 license, which lets users rely on a hosted provider while retaining the option to run its open weights if hosting ends.
Google released Nano Banana 2.1, an image-generation and editing model powered by Gemini 3.6 Flash, with claimed improvements to image quality, text rendering, consistency and editing. Its 1K image price is 3.36 cents—about half the previous model’s price—and it is rolling out across Google products, including Gemini, Search and AI Studio.
OpenAI introduced the Decisions API in public beta, a tool that returns yes/no probabilities, category choices or ratings for text and images, and says it runs about 10 times faster than its Responses API. It currently supports gpt-6-luna at $0.10 per million input tokens, and OpenAI also reduced its paid API tiers to Build, Launch and Grow.
The Technology Innovation Institute in Abu Dhabi introduced Falcon-ASR, a 1.6-billion-parameter speech recognition model focused on Arabic, especially Emirati, that also supports English, French, Spanish and Portuguese. TII reports a 20.92% average word error rate across six Arabic benchmarks—below the best published result it compared against—and 22.73% on its internal Emirati evaluation.
Google DeepMind researchers Sahil Dua and Henrique Schechter Vera announced EmbeddingGemma 2, a 740-million-parameter model that embeds text, code, images, video and audio in a shared space for on-device search and retrieval. Released under the Apache 2.0 license, it supports an 8K-token context and can use as little as 191MB of active RAM for text-only inference; Google says its code benchmark score rose from 68.76 to 78.68 compared with the first EmbeddingGemma.
Mistral has released a public preview of Mistral Large 4, a Europe-trained model with about one trillion parameters; its weights are expected at the end of October. It scored 38 on Artificial Analysis’s Intelligence Index, below Claude Opus 5.5’s 58, while Mistral highlights cybersecurity testing where the model reproduces and patches vulnerabilities—tasks it says rivals Claude and GPT-6 refuse.
Reflection has launched Beam, a text-only, 501-billion-parameter mixture-of-experts model with 23 billion active parameters, trained from scratch for coding, agentic and scientific tasks. The company says training used 23.8 trillion tokens and reports an 80.9 score on SWE-bench Verified; it plans to release the full weights under Apache 2.0 this month.
OpenAI is rolling out GPT-6 with “Intelligent UI,” which can turn answers into interactive charts, forms, buttons and small in-chat tools rather than mostly text. The company says GPT-6 can also answer while it is still reasoning, cutting wait times by 44 percent; the rollout begins with paid plans, followed by free and Go users the next day.
Aleph Alpha has released Kolibri, an open-weight German-English model with 78 billion parameters, about three billion active per token, and a context window of up to one million tokens. The company says it is designed for public administration, aviation, and industry, was trained on GPUs in Germany and Finland, and is available under the Apache 2.0 license.
Google DeepMind introduced Gemini 4 Argon, a frontier AI model for coding, enterprise work, and cybersecurity with state-of-the-art performance on 13 of 19 benchmarks and an industry-leading 1M-token output limit, though it is currently available only to government users and trusted cyber defenders in the Fairwind Program. The model is priced at $4/$20 per million input/output tokens with a 50% introductory discount, and internal tests show it outperforms competitors like GPT-6 Astra and Claude Opus 5.5 on tasks like code migration and software engineering challenges.
Liquid AI released two open-weight decision models: d1-3B, which handles text and images, and experimental d1-omni-600M, which accepts text with images or audio. The company says d1-3B scores 82.9 on its seven-dataset benchmark and answers a question in 16–50 ms on tested NVIDIA Jetson edge devices; d1-omni-600M scores 78.4, with speed figures not yet reported.
NVIDIA says its Nemotron models achieved gold-level results at IMO 2026 and IOI 2026 using fine-tuning and generate-verify-refine inference systems. The IMO system scored 30/42, above the 29-point gold threshold; the IOI system scored 535.4/600, above the gold threshold and the top human score, but its run was unofficial and excluded from the rankings.
The Technology Innovation Institute has released Falcon-Emirati-7B, a 7-billion-parameter model adapted from Falcon-H1-Arabic to understand and generate Emirati Arabic, using native dialect data, cultural material and guided synthetic examples. TII reports scores of 84.83% on its 1,173-question Alyah benchmark and 85.57% on a UAE cultural-understanding test, ahead of the comparison models it evaluated.
Simon Willison compares SVGs generated from the same unusual prompt—an armadillo in fishnet tights jaywalking on Mars—by Claude Opus 5.5, GPT-6.1 Sol, Gemini 3.8 Flash and Mistral Large 4. The post presents the comparison in response to a comment that AI benchmarks have become saturated.
Simon Willison tested whether Claude Opus 5.5 could compose video-game music, asking it to create a simple text-based music format and a playable artifact with example tracks inspired by *The Secret of Monkey Island*. He says the results were surprisingly good and wonders whether competent music composition is a newly emerged capability in text models, though he has not yet compared it with other models.
Simon Willison tested Qwen3.8-27B on a DGX Spark, asking it to add large numbers and spell out the results. With reasoning enabled, the model answered correctly in 167 of 169 one-shot attempts; without reasoning, he ran 30 attempts for each number-size combination.
Benjamin Marie compares Qwen3.8 Flash Next with reasoning off, low, medium and xhigh across 14 benchmarks and 23,331 problems per setting, measuring accuracy and generated tokens. He reports that low is surprisingly competitive with medium, while xhigh uses more than three times as many output tokens as medium and its accuracy gains vary by workload.
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on whether they actually change database records correctly, not just whether they make valid tool calls or sound right. Testing 12 language models on 507 stateful business tasks run 20 times each, they found that 67% of failed attempts still produced clean tool calls with no reported errors—revealing wrong field values, unintended side effects, or incomplete actions. Only three models (GPT-6 Astra, Claude Opus 5.5, Claude Opus 5) retained most of their single-attempt accuracy across repeated trials, while others like Kimi-K3 solved more tasks at least once but failed consistently, and roughly 80% of failures stemmed from tool-handling errors rather than reasoning problems.
Pi released version 1.0 and Pi Durable, with the latter porting the AI agent framework to TypeScript and adding features like crash recovery, concurrent conversations, and hot-swapping of tools while agents run. Google announced Gemini 4 Argon with revised pretraining and post-training methods, while OpenAI released GPT-6.1 Sol as an efficiency-focused update that costs $0.72 per task versus $1.04 for GPT-6 Sol. Black Forest Labs launched FLUX 3 with native 4K image generation, multi-image references, and layout control, along with other AI releases including Upstage's Solar Mini 4 reasoning model and Tavus's Griffin video interaction model.
NVIDIA released Kumo Tabular, an open-source foundation model for tabular data that can predict labels for new rows without training or feature engineering, available in three sizes from 28M to 215M parameters. The model was pretrained entirely on artificial data and ranks first on four benchmarks (TabArena, BeyondArena, TALENT, and ScoringBench) while running 17 times faster than competing systems. Kumo Tabular uses a Transformer architecture with specialized attention mechanisms for columns, rows, and in-context learning to handle classification and regression tasks on structured data.
Benjamin Marie benchmarked 12 quantized GGUF versions of Qwen3.8 Flash Next against the 354 GB BF16 model, measuring accuracy and token efficiency across 42.7 million generated tokens. All 11 standard quantizations—from Q4 to Q1—retained at least 95% of the reference model’s accuracy; a separate pruned-and-quantized Coder release preserved coding performance but lost ground in general knowledge and scientific reasoning.