OpenAI has begun a limited preview of the GPT‑5.6 family: **Sol** (flagship), **Terra** (balanced, 2× cheaper than GPT‑5.5), and **Luna** (fast, low‑cost). Sol introduces a new “max” reasoning mode and an “ultra” mode that uses sub‑agents. Early benchmarks show state‑of‑the‑art performance on coding (Terminal‑Bench 2.1), biology (GeneBench v1), and cybersecurity (ExploitBench, ExploitGym). The models ship with OpenAI’s most robust safety stack to date, but the preview may block or delay some requests.
GLM‑5.2 is Z.ai’s newest open‑source model designed for long‑horizon coding tasks. It supports a solid 1 million‑token context, introduces flexible “effort” levels for balancing speed and capability, and uses the IndexShare architecture to cut per‑token FLOPs by 2.9×. Benchmarks show it outperforms its predecessor (GLM‑5.1) and ranks as the strongest open‑source coding model, closing the gap to leading closed‑source systems.
Z.ai: announced GLM‑5.2, a frontier‑intelligence model with open weights, a 1 M‑token context window and two reasoning effort levels for coding and agentic tasks. OpenCode: reported that GLM‑5.2 has risen to 6th place on their leaderboard within three days of release. ollama: highlighted GLM‑5.2 as the strongest open‑source coding model yet, now available on Ollama’s US cloud powered by NVIDIA AI Blackwell GPUs. GitHub Changelog: warned that Opus 4.6 (fast) will be deprecated in all Copilot experiences on June 29 2026 and urged migration to Opus 4.8 (fast). Cursor: introduced a new /automate skill that lets agents set up automations from plain‑language task descriptions, configuring triggers, instructions and tools automatically.
Kimi.ai: Released the open‑source Kimi‑K2.7‑Code model, reporting +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, +31.5% on MLS Bench Lite, and 30% lower reasoning overthinking. Google Research: Launched Gemini‑SQL2, a text‑to‑SQL system built on Gemini 3.1 Pro that hits state‑of‑the‑art scores on the BIRD benchmark. OpenAI: Added saved Codex rate‑limit resets for Go, Plus, Pro, and Business tiers (one free reset) and a two‑week invite program letting Plus/Pro users earn extra resets by inviting friends. Claude: Made dynamic workflows in Claude Code generally available, letting the model orchestrate parallel sub‑agents for complex tasks like codebase‑wide bug hunts and verify work before returning results.
GitHub announces that Agentic Workflows are now in public preview, offering intelligent automations with guardrails, observability, and cost controls. GitHub Changelog notes the workflows can now use the built‑in GITHUB_TOKEN instead of personal access tokens, improving security and simplicity. OpenCode reports that DeepSeek V4 Pro, Fable 5, and the North Mini Code model (256K context, fully open source) are now available on its platform. OpenRouter launches an Activity explorer that shows real‑time spending, token usage, cache hit rates, agents, and trends for models like Fable. RyanLee shares that his high‑performance MSA kernel library is open‑source and that the M3 weights are expected to be released on Friday, with a link to the GitHub paper.
Claude Fable 5 is Anthropic’s latest Mythos‑class model, released with safety classifiers that route risky queries to Claude Opus 4.8. It claims state‑of‑the‑art performance on coding benchmarks, dramatically faster codebase migrations, and higher token efficiency than previous Claude models.
Google announced Gemini 3.5 Live Translate, an audio model that streams speech‑to‑speech translation in real time for more than 70 languages. It is available now in public preview via the Gemini Live API, in private preview for Google Meet, and globally in the Google Translate mobile app.
Claude announced Claude Fable 5, a Mythos‑class model deemed safe for general use and now available everywhere, while Claude Mythos 5 stays limited to Glasswing partners. OpenRouter, Cursor, Visual Studio Code and Devin Desktop all reported that Claude Fable 5 is now live on their platforms, with Cursor noting a 72.9% score on CursorBench, 8 points above the previous best. Google DeepMind highlighted its 3.5 Live Translate, which streams speech into over 70 languages while preserving tone, pace and pitch for natural conversation. Kimi.ai introduced Kimi Work, a local desktop AI agent that can run up to 300 parallel agents and, via a WebBridge extension, navigate browsers to search, scroll and interact.
NVIDIA’s Nemotron‑3 Ultra is a frontier‑scale LLM with 550 B total (55 B active) parameters, a hybrid LatentMixture‑of‑Experts (LatentMoE) architecture, and up to 1 M token context length. It runs on NVIDIA GPUs (minimum 4 × B200/H100) and offers configurable reasoning traces, multi‑token speculative decoding, and multilingual support. Benchmarks show strong performance on agentic, reasoning, and long‑context tasks.
Cursor’s latest canvas update introduces **Design Mode**, letting users point‑and‑click UI elements for edits, and a **Context Usage Report** that visualizes token distribution. Additional quality‑of‑life tweaks include full‑screen shared canvases, embedded action buttons, better error handling, and richer chart styling.
NVIDIA AI announced the release of Nemotron 3 Ultra, a 550 B MoE model that speeds inference fivefold, lowers agentic task costs up to 30 % and excels at coding, deep research, and long‑horizon planning. OpenCode noted that Nemotron 3 Ultra is now free with 1 M context and fully open source. Ollama said the model is available on its cloud platform, offering launch commands for Claude, Hermes and OpenClaw. OpenAI introduced a new memory system for ChatGPT that automatically tracks important details, doubles memory capacity for Plus and Pro users in the US, and lets users review and steer remembered content via a summary. Cursor added an interactive context‑usage report in its canvas, breaking down token distribution across prompts, tools, rules and skills.
NVIDIA and Microsoft announced a full‑stack solution for building and running AI agents on Windows PCs, enterprise workstations, and Azure. New hardware (RTX Spark laptops, DGX Station for Windows, RTX PRO 6000 Blackwell servers) pairs with NVIDIA OpenShell runtime, open‑source models on Microsoft Foundry, and GPU‑accelerated Microsoft Fabric. The goal is to let developers code, tune, and deploy long‑running, secure agents locally or in the cloud.
Google Gemma: Announces Gemma 4 12B, a unified encoder‑free multimodal model for laptops released under an Apache 2.0 license. Google AI Developers: Highlights that Gemma 4 12B bridges their mobile E4B and larger 26B MoE models, offering frontier‑class reasoning and native audio. NVIDIA: Notes local AI agents advancing on DGX Spark and RTX PCs, with OpenShell arriving on Windows, new agentic AI optimizations, Broadcast 2.2, and upcoming RTX acceleration for Adobe apps and Blender. Visual Studio Code: Reports May updates—Agents window now stable, BYOK with air‑gapped support, and an integrated browser that can emulate devices and preview HTML without extensions. Ideogram: Introduces Ideogram 4.0, an open image model with downloadable weights, fine‑tuning on personal data, and availability across all plans and the API.
NVIDIA unveiled RTX Spark, a new class of Windows PCs built for on‑device AI agents. With up to 1 petaflop of AI compute, 128 GB of unified memory, and new security primitives (OpenShell), the platform promises faster, private inference for popular open‑source agents such as Hermes and OpenClaw. Similar capabilities are extended to Linux via DGX Spark, while multi‑GPU optimizations boost llama.cpp and ComfyUI performance.
Gemma 4 12B is Google’s new 12‑billion‑parameter multimodal model that runs locally on consumer laptops (≈16 GB VRAM). It eliminates separate vision and audio encoders, delivers reasoning close to the larger 26 B Mixture‑of‑Experts model, and is released under an Apache 2.0 license with full tool‑chain support.
Qwen 3.7‑Plus is a new multimodal agent model that adds vision capabilities to the strong text backbone of Qwen 3.7. It can read screens, interact with GUIs, and generate code from visual references while keeping the coding and tool‑use strengths of its predecessor. Benchmarks show notable gains in several coding‑related tasks, especially in terminal‑based and spreadsheet benchmarks.
OpenAI’s latest frontier models, including GPT‑5.5, and the Codex coding assistant are now generally available through Amazon Bedrock on AWS. This integration lets enterprises use OpenAI’s most advanced AI capabilities within their existing AWS security, compliance, and deployment workflows.
Qwen: introduces Qwen3.7-Plus, a multimodal agent model that unifies vision and language with both GUI and CLI operation and serves as a coding and productivity assistant. OpenAI: frontier models and Codex are now generally available on AWS via Amazon Bedrock, extending enterprise security, compliance, and governance workflows. xAI: Composer 2.5 is now inside Grok Build, described as a fast, highly intelligent model for long‑running tasks and complex instructions. LangChain: highlights Fleet for secure agent access to private resources and adds LangSmith LLM Gateway spend limits that return a 402 error when caps are hit. Google Antigravity: is becoming a scientific workbench with a Science Skills bundle that runs complex workflows like protein analysis using Alpha* models and dozens of databases; Google Gemma: releases the first gemma‑skills iteration, enabling agents to build with Gemma, use MTP for speed, pick model size, and locate up‑to‑date resources. ClaudeDevs: resets 5‑hour and weekly rate limits for Pro/Max plans and fixes excessive parallel subagents; Cursor: raises usage limits for Teams and adds a Premium seat with 5× usage at 3× cost; Visual Studio Code: demos orchestrating agents via the VS Code Agents window; NVIDIA: adds real‑time AI media tools including Synthetic Video Detector (up to 92% accuracy, 22 ms latency), RTX Video Super Resolution and Frame Generation; Vercel: enables remote execution of Conductor’s parallel coding agents on fast Sandboxes; Perplexity: launches Search as Code, a new architecture that writes Python to call its search stack directly, now default in the Perplexity Agent API.
MiniMax released its latest M‑series model, **MiniMax‑M3**, on June 1 2026. The model is marketed for agentic reasoning, tool use, coding, multimodal chat input, and long‑context tasks. It follows a series of MiniMax models (M2.5, M2.1) that already claimed state‑of‑the‑art (SOTA) performance in programming, code refactoring, and tool calling.
LangChain – Introduces Managed Deep Agents that retain the familiar project layout (AGENTS.md, skills/, subagents/, tools.json) and adds a Context Hub for persisting and updating agent context across sessions; a technical roundtable on June 17 in Munich will cover production‑grade agents, agent harnesses, and the Deep Agents SDK. MiniMax – Launches MiniMax M3, an open‑weights model that merges a 1 million‑token context window, frontier coding, agentic abilities, and native multimodal (image/video) support, with benchmark scores highlighted and a 50 % discount for the first week; the model appears automatically in the Hermes Agent picker and credits the Teknium and Nous teams. OpenRouter – Announces that MiniMax‑M3 is now available on its platform, offering the same 1 M‑token context, frontier coding, agentic performance, and multimodal capabilities. Visual Studio Code – Promotes its new VS Code Learn Series episode “Extending Agents,” which teaches how to use tools, agent plugins, and third‑party agents within the editor.
OpenRouter (@OpenRouter) has integrated its models into ComfyUI workflows, allowing users to leverage OpenRouter models directly within ComfyUI. GitHub (@github) highlights the 2026 Partner Pack, offering exclusive discounts and perks for maintainers. The GitHub Innovation Graph provides economic data on trends in GDP, inequality, and emissions, which researchers find valuable. Google AI Developers (@googleaidevs) showcases successful implementations of Managed Agents in the Gemini API, including Eigent_AI's root cause analysis and llama_index's document processing template. NVIDIA (@nvidia) discusses the Dell AI Factory with NVIDIA, which enables companies to build, run, and scale AI, with NemoClaw powering agentic AI on-prem. OpenAI (@OpenAI) shares Terence Tao's experience with AI, which gives researchers more freedom to experiment and pursue unconventional ideas. LangChain (@LangChain) emphasizes the efficiency of LangSmith Sandboxes, which pause automatically when idle, and encourages users to create agents using everyday language with LangSmith Fleet.
Claude Code has released version 2.1.152, which includes several new features and bug fixes. The update applies review findings to the working tree after a review, simplifies code, and improves the user experience.
Claude Opus 4.8, Anthropic's latest Opus model, is now available in GitHub Copilot. This model demonstrates a clear step forward in code understanding and generation across various real-world coding tasks.
OpenAI and Thrive Holdings collaborated to build Tax AI, a self-improving tax agent that automates tax preparation and improves over time. The system uses Codex to turn production use into structured signals that fuel autonomous improvement.
OpenRouter now supports "apply_patch," a server tool that lets models propose file edits using V4A diffs through the Responses API. The model generates a patch, and OpenRouter validates the diff syntax server-side. This feature allows for more efficient and accurate file editing. xAI has released grok-build-0.1 in public beta via the xAI API. This model powers the Grok Build CLI and excels at agentic coding, priced at $1/m input and $2/m output. Google AI has released an episode of Release Notes featuring the architects of Gemini, including @JeffDean, @koraykv, @OriolVinyalsML, and @NoamShazeer. They discuss their journey and the people behind the model. LangChain has released LangSmith LLM Gateway, which enforces spend limits and redacts PII before requests reach the model. They also announced Deep Agents v0.6, which makes harness profiles a first-class abstraction, allowing for production-grade performance at lower costs. NVIDIA has announced a new era of PC, but the details are unclear. OpenAI has launched Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities. They are also expanding trusted access to GPT-Rosalind for select U.S. government and allied partners.
MiniMax has released a new AI agent model, MiniMax M2.7, which matches top closed models in benchmarks. However, the model's license has been changed to restrict commercial use, sparking controversy among developers.
LangChain (@LangChain) LangChain has released a new course on LangSmith Fleet Essentials, allowing users to build, use, and manage agent fleets for complex tasks without coding. The course is a quickstart guide to building and improving agents. LangSmith Engine is also mentioned, which optimizes self-improving loops. A keynote from @hwchase17 highlighted the future of agents. Google DeepMind (@GoogleDeepMind) Google DeepMind's Gemini for Science tools aim to help scientists achieve their next breakthrough.
GitHub Spark is a platform for building intelligent apps with AI capabilities. It allows users to create full-stack applications using natural language, visual tools, or code, with instant previews and one-click deployment.
GitHub is hosting a Maintainer AMA in their community on May 27 from 8 a.m. to 1 p.m. PT, where contributors can ask questions and show appreciation for open-source projects like OpenClaw and Kubernetes. LangChain is using traces to build evaluations for production agents and has developed Mission Control, a decoupled application for deploying, configuring, and troubleshooting self-hosted LangChain infrastructure within Kubernetes. xAI has reset Grok Build usage limits for all accounts after receiving feedback on the Grok Build Beta, with the team continuing to collect feedback to improve the service. Anthropic has published a blog post discussing the importance of evolving access and permissions for agents as their capabilities grow, using sandboxing to limit potentially destructive actions in their own products.