Model Signal logo Model Signal Fast, verified AI updates

Tag

API

Published stories tagged with API.

AI Models 3 min read

Latest from X - 2026-06-30 to 2026-07-02

Google Research: Introduced TabFM, a foundation model for tabular data classification and regression that can generate high‑quality predictions on unseen tables in a single forward pass. ClaudeDevs: Raised Claude Platform API rate limits for all users, removed spend‑based tiers, and gave the latest Sonnet and Haiku models five‑times higher limits at the top tier. Google AI: Launched two workflow updates - one model for ultra‑fast image generation and another to instantly animate those images, both at a fraction of the usual cost, highlighted by Nano Banana 2 Lite. GitHub Changelog: Announced that Gemini 2.5 Pro and Gemini 3 Flash will be deprecated in all Copilot experiences on July 31 2026, urging migration to Gemini 3.1 Pro or Gemini 3.5 Flash beforehand. Z.ai: Rolled out ZCode, the official development environment for GLM‑5.2, offering 1.5× usage quota for GLM Coding Plan subscribers, BYOK support, and cross‑platform availability on macOS, Windows, and Linux. GitHub: Made Kimi Moonshot’s Kimi K2.7 Code generally available in GitHub Copilot as the first open‑weight model selectable in the model picker, noting lower cost with performance comparable to top‑tier models.

AI Models 4 min read

Mistral OCR 4 Brings Structured Document Extraction with Bounding Boxes and Multilingual Support

Mistral OCR 4 is a compact, self‑hostable OCR model that returns not only extracted text but also bounding boxes, block types (titles, tables, equations, etc.) and per‑word confidence scores. It supports 170 languages, runs in a single container, and is available via API (or Document AI layer). Independent human evaluations show a 72 % win‑rate over leading OCR systems, and benchmark scores place it at the top of public OCR suites while delivering up to 8× lower cost and 17× lower latency for high‑volume use cases.

AI Models 4 min read

Previewing GPT‑5.6 Sol, Terra, and Luna

OpenAI has begun a limited preview of the GPT‑5.6 family: **Sol** (flagship), **Terra** (balanced, 2× cheaper than GPT‑5.5), and **Luna** (fast, low‑cost). Sol introduces a new “max” reasoning mode and an “ultra” mode that uses sub‑agents. Early benchmarks show state‑of‑the‑art performance on coding (Terminal‑Bench 2.1), biology (GeneBench v1), and cybersecurity (ExploitBench, ExploitGym). The models ship with OpenAI’s most robust safety stack to date, but the preview may block or delay some requests.

AI Models 2 min read

Latest from X - 2026-06-26

OpenAI: OpenAI announced a limited preview of the GPT‑5.6 family—Sol, Terra and Luna. Sol is the new flagship, described as a step‑function improvement over GPT‑5.5. Terra delivers performance competitive with GPT‑5.5 at roughly 2× lower cost, and Luna is the most cost‑efficient model, offering strong capability at the lowest price. The company plans to make the three models generally available in the coming weeks, but the preview is currently limited to a small group of trusted partners in Codex and the API at the request of the U.S. government.

AI Models 4 min read

GLM‑5.2 Brings 1M‑Token Context to Coding Agents

GLM‑5.2 is Z.ai’s newest open‑source model designed for long‑horizon coding tasks. It supports a solid 1 million‑token context, introduces flexible “effort” levels for balancing speed and capability, and uses the IndexShare architecture to cut per‑token FLOPs by 2.9×. Benchmarks show it outperforms its predecessor (GLM‑5.1) and ranks as the strongest open‑source coding model, closing the gap to leading closed‑source systems.

Coding 2 min read

Latest from X - 2026-06-12 to 2026-06-13

OpenCode: Kimi 2.7 Code is now available in Go with image support, optimized for coding and priced similarly to 2.6. Kimi.ai: Builders using the Kimi K2.7 Code API can earn 20%–30% extra quota by topping up $100+ before July 2, with one bonus per account. ollama: The Kimi‑K2.7‑Code model is now hosted on Ollama’s US cloud on NVIDIA B300 GPUs, keeping data private and never used for training; try it with “ollama launch claude --model”. Z.ai: GLM‑5.2, the new flagship model, is now accessible to all GLM Coding Plan users—including Lite, Pro, Max, and Team tiers.

Coding 4 min read

NVIDIA AI announced the release of NVIDIA Nemotron‑3 Ultra 550B (55B active) – What Developers Need to Know

NVIDIA’s Nemotron‑3 Ultra is a frontier‑scale LLM with 550 B total (55 B active) parameters, a hybrid LatentMixture‑of‑Experts (LatentMoE) architecture, and up to 1 M token context length. It runs on NVIDIA GPUs (minimum 4 × B200/H100) and offers configurable reasoning traces, multi‑token speculative decoding, and multilingual support. Benchmarks show strong performance on agentic, reasoning, and long‑context tasks.

Coding 3 min read

Latest from X - 2026-06-03

Google Gemma: Announces Gemma 4 12B, a unified encoder‑free multimodal model for laptops released under an Apache 2.0 license. Google AI Developers: Highlights that Gemma 4 12B bridges their mobile E4B and larger 26B MoE models, offering frontier‑class reasoning and native audio. NVIDIA: Notes local AI agents advancing on DGX Spark and RTX PCs, with OpenShell arriving on Windows, new agentic AI optimizations, Broadcast 2.2, and upcoming RTX acceleration for Adobe apps and Blender. Visual Studio Code: Reports May updates—Agents window now stable, BYOK with air‑gapped support, and an integrated browser that can emulate devices and preview HTML without extensions. Ideogram: Introduces Ideogram 4.0, an open image model with downloadable weights, fine‑tuning on personal data, and availability across all plans and the API.

Coding 3 min read

Qwen 3.7‑Plus: Multimodal Coding Agent with Vision‑Language Upgrade

Qwen 3.7‑Plus is a new multimodal agent model that adds vision capabilities to the strong text backbone of Qwen 3.7. It can read screens, interact with GUIs, and generate code from visual references while keeping the coding and tool‑use strengths of its predecessor. Benchmarks show notable gains in several coding‑related tasks, especially in terminal‑based and spreadsheet benchmarks.

AI Models 6 min read

Latest from X - 2026-06-01 to 2026-06-02

Qwen: introduces Qwen3.7-Plus, a multimodal agent model that unifies vision and language with both GUI and CLI operation and serves as a coding and productivity assistant. OpenAI: frontier models and Codex are now generally available on AWS via Amazon Bedrock, extending enterprise security, compliance, and governance workflows. xAI: Composer 2.5 is now inside Grok Build, described as a fast, highly intelligent model for long‑running tasks and complex instructions. LangChain: highlights Fleet for secure agent access to private resources and adds LangSmith LLM Gateway spend limits that return a 402 error when caps are hit. Google Antigravity: is becoming a scientific workbench with a Science Skills bundle that runs complex workflows like protein analysis using Alpha* models and dozens of databases; Google Gemma: releases the first gemma‑skills iteration, enabling agents to build with Gemma, use MTP for speed, pick model size, and locate up‑to‑date resources. ClaudeDevs: resets 5‑hour and weekly rate limits for Pro/Max plans and fixes excessive parallel subagents; Cursor: raises usage limits for Teams and adds a Premium seat with 5× usage at 3× cost; Visual Studio Code: demos orchestrating agents via the VS Code Agents window; NVIDIA: adds real‑time AI media tools including Synthetic Video Detector (up to 92% accuracy, 22 ms latency), RTX Video Super Resolution and Frame Generation; Vercel: enables remote execution of Conductor’s parallel coding agents on fast Sandboxes; Perplexity: launches Search as Code, a new architecture that writes Python to call its search stack directly, now default in the Perplexity Agent API.

Coding 3 min read

Latest from X - 2026-06-01

LangChain – Introduces Managed Deep Agents that retain the familiar project layout (AGENTS.md, skills/, subagents/, tools.json) and adds a Context Hub for persisting and updating agent context across sessions; a technical roundtable on June 17 in Munich will cover production‑grade agents, agent harnesses, and the Deep Agents SDK. MiniMax – Launches MiniMax M3, an open‑weights model that merges a 1 million‑token context window, frontier coding, agentic abilities, and native multimodal (image/video) support, with benchmark scores highlighted and a 50 % discount for the first week; the model appears automatically in the Hermes Agent picker and credits the Teknium and Nous teams. OpenRouter – Announces that MiniMax‑M3 is now available on its platform, offering the same 1 M‑token context, frontier coding, agentic performance, and multimodal capabilities. Visual Studio Code – Promotes its new VS Code Learn Series episode “Extending Agents,” which teaches how to use tools, agent plugins, and third‑party agents within the editor.

Coding 1 min read

Latest from X - 2026-05-31 (VS Code)

Visual Studio Code (@code): The team introduced Chronicle, an experimental feature that records Copilot chat sessions in a local SQLite database. This logging aims to give developers insight into their AI-assisted coding workflow. They suggest it can help boost productivity.

Coding 1 min read

Latest from X - 2026-05-29 (ClaudeDevs)

ClaudeDevs (@ClaudeDevs) announced an update to Opus 4.8, which now allows for the addition of system instructions mid-conversation without breaking the prompt cache. This improvement is expected to reduce the cost and latency of API requests. The update aims to enhance the efficiency of the system.

AI Models 3 min read

Latest from X - 2026-05-30

OpenRouter (@OpenRouter) has integrated its models into ComfyUI workflows, allowing users to leverage OpenRouter models directly within ComfyUI. GitHub (@github) highlights the 2026 Partner Pack, offering exclusive discounts and perks for maintainers. The GitHub Innovation Graph provides economic data on trends in GDP, inequality, and emissions, which researchers find valuable. Google AI Developers (@googleaidevs) showcases successful implementations of Managed Agents in the Gemini API, including Eigent_AI's root cause analysis and llama_index's document processing template. NVIDIA (@nvidia) discusses the Dell AI Factory with NVIDIA, which enables companies to build, run, and scale AI, with NemoClaw powering agentic AI on-prem. OpenAI (@OpenAI) shares Terence Tao's experience with AI, which gives researchers more freedom to experiment and pursue unconventional ideas. LangChain (@LangChain) emphasizes the efficiency of LangSmith Sandboxes, which pause automatically when idle, and encourages users to create agents using everyday language with LangSmith Fleet.

AI Models 2 min read

Claude Code Releases

Claude Code has released version 2.1.152, which includes several new features and bug fixes. The update applies review findings to the working tree after a review, simplifies code, and improves the user experience.

AI Models 4 min read

Latest from X - 2026-05-29

OpenRouter now supports "apply_patch," a server tool that lets models propose file edits using V4A diffs through the Responses API. The model generates a patch, and OpenRouter validates the diff syntax server-side. This feature allows for more efficient and accurate file editing. xAI has released grok-build-0.1 in public beta via the xAI API. This model powers the Grok Build CLI and excels at agentic coding, priced at $1/m input and $2/m output. Google AI has released an episode of Release Notes featuring the architects of Gemini, including @JeffDean, @koraykv, @OriolVinyalsML, and @NoamShazeer. They discuss their journey and the people behind the model. LangChain has released LangSmith LLM Gateway, which enforces spend limits and redacts PII before requests reach the model. They also announced Deep Agents v0.6, which makes harness profiles a first-class abstraction, allowing for production-grade performance at lower costs. NVIDIA has announced a new era of PC, but the details are unclear. OpenAI has launched Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities. They are also expanding trusted access to GPT-Rosalind for select U.S. government and allied partners.