Jump to channels Log in

AI Model Releases

1 member · about 12 posts a day

New AI models as they ship, from frontier LLMs and open-weight releases to image, video and audio models, with technical analysis and independent benchmarks.

28 stories this week
0

[AINews] not much happened today latent.space

OpenAI safety researchers Tomek Korbak, Mikita Balesni and Jasmine Wang say they were fired after raising concerns about AI safety and the company’s work with evaluator METR; OpenAI reportedly says they mishandled confidential information. The roundup also reports releases including OpenAI’s GPT-6.1 Sol Ultrafast and Anthropic’s Claude Haiku 5.5, alongside new research and benchmark findings on agent safety, model performance and infrastructure.

0

The model that didn't exist, so you made it yourself huggingface.co

Hugging Face author Yvrj Sharma says its ML-intern agent helped build six models and LoRAs over several days, with projects costing about $103 in total compute. Examples include a 0.8B image-prompt rewriter that runs on a CPU and achieved 99.7% valid outputs, a citrus-disease vision model that raised accuracy from 14.9% to 52.8% on 335 test images, and a four-step image generator that improved GenEval from 0.509 to 0.536. The agent plans and runs training, evaluation and publication jobs after getting a budget approved.

0

Introducing Mistral Large 4: Le chonk simonwillison.net

Mistral has released a preview of Mistral Large 4, a 1-trillion-parameter model with 49 billion active parameters, trained on 3,800 NVIDIA Grace Blackwell GPUs. The API preview offers “none” and “high” reasoning modes; Mistral says open weights will arrive at the end of the month, and the model scores 38 on Artificial Analysis, up from 9 for Mistral Large 3.

1

Introducing Claude Haiku 5.5 simonwillison.net Breaking

Anthropic released Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with rates five times higher above that; blogger Simon Willison notes its tokenizer uses about 25% more tokens than Haiku 4.5. Anthropic also halved Sonnet 5.5 cache-read prices and announced monthly API credits of up to $100 for Max 5x, $200 for Max 20x, and $500 pooled for Team subscribers; credits expire monthly.

1

[AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics latent.space

OpenAI released 722 manuscripts from an unreleased internal math model, covering 372 groups of results drawn from an evaluation of about 4,000 problems. The company says the work used roughly three hours of ChatGPT Pro compute per result; commentators highlight a claimed quasi-Riemann hypothesis result, but the manuscripts and proofs have not been independently verified.

0

EmbeddingGemma 2 simonwillison.net

Simon Willison argues that embedding models should be available under open licenses because replacing a discontinued proprietary model can force users to recalculate millions of stored vectors. He welcomes EmbeddingGemma 2’s Apache 2.0 license, which lets users rely on a hosted provider while retaining the option to run its open weights if hosting ends.

1

Google's new image model Nano Banana 2.1 generates better images for less money the-decoder.com

Google released Nano Banana 2.1, an image-generation and editing model powered by Gemini 3.6 Flash, with claimed improvements to image quality, text rendering, consistency and editing. Its 1K image price is 3.36 cents—about half the previous model’s price—and it is rolling out across Google products, including Gemini, Search and AI Studio.

0

OpenAI launches Decisions API that reduces complex evaluations to yes, no, or pick one the-decoder.com

OpenAI introduced the Decisions API in public beta, a tool that returns yes/no probabilities, category choices or ratings for text and images, and says it runs about 10 times faster than its Responses API. It currently supports gpt-6-luna at $0.10 per million input tokens, and OpenAI also reduced its paid API tiers to Build, Launch and Grow.

0

Introducing Falcon ASR huggingface.co

The Technology Innovation Institute in Abu Dhabi introduced Falcon-ASR, a 1.6-billion-parameter speech recognition model focused on Arabic, especially Emirati, that also supports English, French, Spanish and Portuguese. TII reports a 20.92% average word error rate across six Arabic benchmarks—below the best published result it compared against—and 22.73% on its internal Emirati evaluation.

1
posted by matt2000 2 days ago

EmbeddingGemma 2: an open, lightweight multimodal embedding model blog.google

Google DeepMind researchers Sahil Dua and Henrique Schechter Vera announced EmbeddingGemma 2, a 740-million-parameter model that embeds text, code, images, video and audio in a shared space for on-device search and retrieval. Released under the Apache 2.0 license, it supports an 8K-token context and can use as little as 191MB of active RAM for text-only inference; Google says its code benchmark score rose from 68.76 to 78.68 compared with the first EmbeddingGemma.

1

Mistral Large 4 is Europe's trillion-parameter answer to US models that refuse security work the-decoder.com

Mistral has released a public preview of Mistral Large 4, a Europe-trained model with about one trillion parameters; its weights are expected at the end of October. It scored 38 on Artificial Analysis’s Intelligence Index, below Claude Opus 5.5’s 58, while Mistral highlights cybersecurity testing where the model reproduces and patches vulnerabilities—tasks it says rivals Claude and GPT-6 refuse.

1

[AINews] Reflection Beam - 501B-A23B American Open Model latent.space

Reflection has launched Beam, a text-only, 501-billion-parameter mixture-of-experts model with 23 billion active parameters, trained from scratch for coding, agentic and scientific tasks. The company says training used 23.8 trillion tokens and reports an 80.9 score on SWE-bench Verified; it plans to release the full weights under Apache 2.0 this month.

0

ChatGPT with GPT-6 ditches mostly text output for interactive UI with charts, buttons, and mini apps the-decoder.com

OpenAI is rolling out GPT-6 with “Intelligent UI,” which can turn answers into interactive charts, forms, buttons and small in-chat tools rather than mostly text. The company says GPT-6 can also answer while it is still reasoning, cutting wait times by 44 percent; the rollout begins with paid plans, followed by free and Go users the next day.

1

Aleph Alpha releases Kolibri, an open-weight model that makes the case for European AI sovereignty the-decoder.com

Aleph Alpha has released Kolibri, an open-weight German-English model with 78 billion parameters, about three billion active per token, and a context window of up to one million tokens. The company says it is designed for public administration, aviation, and industry, was trained on GPUs in Germany and Finland, and is available under the Apache 2.0 license.

1

[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output latent.space

Google DeepMind introduced Gemini 4 Argon, a frontier AI model for coding, enterprise work, and cybersecurity with state-of-the-art performance on 13 of 19 benchmarks and an industry-leading 1M-token output limit, though it is currently available only to government users and trusted cyber defenders in the Fairwind Program. The model is priced at $4/$20 per million input/output tokens with a 50% introductory discount, and internal tests show it outperforms competitors like GPT-6 Astra and Claude Opus 5.5 on tasks like code migration and software engineering challenges.

0

Multimodal open d1 decision models for the edge huggingface.co

Liquid AI released two open-weight decision models: d1-3B, which handles text and images, and experimental d1-omni-600M, which accepts text with images or audio. The company says d1-3B scores 82.9 on its seven-dataset benchmark and answers a question in 16–50 ms on tested NVIDIA Jetson edge devices; d1-omni-600M scores 78.4, with speed figures not yet reported.

0

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO huggingface.co

NVIDIA says its Nemotron models achieved gold-level results at IMO 2026 and IOI 2026 using fine-tuning and generate-verify-refine inference systems. The IMO system scored 30/42, above the 29-point gold threshold; the IOI system scored 535.4/600, above the gold threshold and the top human score, but its run was unofficial and excluded from the rankings.

0

Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance huggingface.co

The Technology Innovation Institute has released Falcon-Emirati-7B, a 7-billion-parameter model adapted from Falcon-H1-Arabic to understand and generate Emirati Arabic, using native dialect data, cultural material and guided synthetic examples. TII reports scores of 84.83% on its 1,173-question Alyah benchmark and 85.57% on a UAE cultural-understanding test, ahead of the comparison models it evaluated.

0

Mistral Large 4 simonwillison.net

Simon Willison compares SVGs generated from the same unusual prompt—an armadillo in fishnet tights jaywalking on Mars—by Claude Opus 5.5, GPT-6.1 Sol, Gemini 3.8 Flash and Mistral Large 4. The post presents the comparison in response to a comment that AI benchmarks have become saturated.

0

Scrimshaw Jukebox simonwillison.net Wildcard

Simon Willison tested whether Claude Opus 5.5 could compose video-game music, asking it to create a simple text-based music format and a playable artifact with example tracks inspired by *The Secret of Monkey Island*. He says the results were surprisingly good and wonders whether competent music composition is a newly emerged capability in text models, though he has not yet compared it with other models.

0

Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh kaitchup.substack.com

Benjamin Marie compares Qwen3.8 Flash Next with reasoning off, low, medium and xhigh across 14 benchmarks and 23,331 problems per setting, measuring accuracy and generated tokens. He reports that low is surprisingly competitive with medium, while xhigh uses more than three times as many output tokens as medium and its accuracy gains vary by workload.

0

The Agent Said It Was Done. The Database Disagreed. huggingface.co Wildcard

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on whether they actually change database records correctly, not just whether they make valid tool calls or sound right. Testing 12 language models on 507 stateful business tasks run 20 times each, they found that 67% of failed attempts still produced clean tool calls with no reported errors—revealing wrong field values, unintended side effects, or incomplete actions. Only three models (GPT-6 Astra, Claude Opus 5.5, Claude Opus 5) retained most of their single-attempt accuracy across repeated trials, while others like Kimi-K3 solved more tasks at least once but failed consistently, and roughly 80% of failures stemmed from tool-handling errors rather than reasoning problems.

0

[AINews] Pi 1.0, Pi Durable, and AIE NYC latent.space

Pi released version 1.0 and Pi Durable, with the latter porting the AI agent framework to TypeScript and adding features like crash recovery, concurrent conversations, and hot-swapping of tools while agents run. Google announced Gemini 4 Argon with revised pretraining and post-training methods, while OpenAI released GPT-6.1 Sol as an efficiency-focused update that costs $0.72 per task versus $1.04 for GPT-6 Sol. Black Forest Labs launched FLUX 3 with native 4K image generation, multi-image references, and layout control, along with other AI releases including Upstage's Solar Mini 4 reasoning model and Tavus's Griffin video interaction model.

0

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction huggingface.co

NVIDIA released Kumo Tabular, an open-source foundation model for tabular data that can predict labels for new rows without training or feature engineering, available in three sizes from 28M to 215M parameters. The model was pretrained entirely on artificial data and ranks first on four benchmarks (TabArena, BeyondArena, TALENT, and ScoringBench) while running 17 times faster than competing systems. Kumo Tabular uses a Transformer architecture with specialized attention mechanisms for columns, rows, and in-context learning to handle classification and regression tasks on structured data.

0

Qwen3.8 Flash Next GGUF Benchmark: Q4 to Q1 Accuracy and Token Efficiency kaitchup.substack.com

Benjamin Marie benchmarked 12 quantized GGUF versions of Qwen3.8 Flash Next against the 354 GB BF16 model, measuring accuracy and token efficiency across 42.7 million generated tokens. All 11 standard quantizations—from Q4 to Q1—retained at least 95% of the reference model’s accuracy; a separate pruned-and-quantized Coder release preserved coding performance but lost ground in general knowledge and scientific reasoning.