2026-07-24
Voice now coordinates Codex work
OpenAI's current Voice settings split the experience three ways: Live is the newest, simultaneous listen-and-speak option; Advanced is the earlier real-time mode with supported…
Last verified 2026-07-24 · Primary source
2026-07-16
Kimi K3 brings a 2.8T open-weight model into the frontier conversation
Moonshot launched Kimi K3 on July 16, and the interesting part is not just the 2.8 trillion parameters. K3 is a sparse MoE model — it activates 16 of 896 experts per token — with…
Last verified 2026-07-22 · Primary source
2026-07-09
ChatGPT Work turns the assistant into a standing agent
OpenAI unveiled ChatGPT Work on July 9 — an agent, powered by GPT-5.6, that can stay with a task for hours rather than answering one prompt at a time. It pulls context from…
Last verified 2026-07-22 · Primary source
2026-07-09
GPT-5.6 splits into three durable tiers — and Sol games its own eval
OpenAI's naming scheme changes with GPT-5.6: the number tracks the model generation, but Sol , Terra , and Luna are capability tiers that can now advance on their own cadence…
Last verified 2026-07-22 · Primary source
2026-07-08
GPT-Live rebuilds voice mode as full-duplex
OpenAI shipped GPT-Live on July 8 — GPT-Live-1 for Go, Plus, and Pro subscribers, and GPT-Live-1 mini for free users — and the headline change is architectural, not cosmetic.…
Last verified 2026-07-22 · Primary source
2026-07-06
AppLess asks whether the phone still needs apps
Rabi Shanker Guha, CEO of Thesys, posted the clearest version yet of the generative-interface thesis: "Imagine you never needed an app again." Every phone screen is generated…
Last verified 2026-07-22 · Primary source
2026-07-01
Claude Fable 5 is available again
Anthropic redeployed Claude Fable 5 globally on July 1, after the U.S. government lifted the export controls that had suspended it on June 12. It is back in Claude, Claude Code,…
Last verified 2026-07-22 · Primary source
2026-07-01
Sakana Fugu hides a multi-agent system behind one model endpoint
Sakana AI's Fugu is a different answer to model selection: don't pick one. It exposes a pool of frontier models through a single OpenAI-compatible API, then learns which agents…
Last verified 2026-07-22 · Primary source
2026-06-30
Claude Sonnet 5 turns the balanced tier into an agent
Anthropic announced Claude Sonnet 5 on June 30. The short version: the Sonnet tier now does work that needed an Opus model a few months ago. It can plan, use browsers and…
Last verified 2026-07-22 · Primary source
2026-06-29
Cursor puts the agent queue on iPhone
Cursor's iOS app is now available in public beta for paid plans. The pitch is not "write Swift on a phone"; it's "keep the agent queue moving when you're away from the laptop."…
Last verified 2026-07-22 · Primary source
2026-06-27
Fable 5 is the model launch that became a policy story
Anthropic launched Claude Fable 5 and Claude Mythos 5 on June 9 as the first public Mythos-class release: same core capability as Mythos 5, but with stricter safeguards that…
Last verified 2026-07-22 · Primary source
2026-06-09
Claude Fable 5 briefly put a Mythos-class model in everyone's hands
Anthropic shipped Claude Fable 5 on June 9 — the first Mythos-class model to go generally available before the access suspension noted above. The Mythos Preview and Glasswing…
Last verified 2026-07-22 · Primary source
2026-06-04
Anthropic puts a number on "AI building AI"
The Karpathy hire below was the staffing thesis. This is the measurement. Anthropic's new research arm, the Anthropic Institute, published When AI builds itself — a look at how…
Last verified 2026-07-22 · Primary source
2026-06-03
Codex expands beyond developers — plugins, Sites, and annotations
OpenAI announced that more than 5 million people now use Codex every week , with non-developers — analysts, marketers, operators, designers, researchers, investors, bankers —…
Last verified 2026-07-22 · Primary source
2026-06-01
Cursor's Auto Review run mode: longer runs, fewer approval prompts
Cursor shipped Auto Review — a new run mode that lets the agent work for longer stretches with fewer approval prompts while keeping execution safe. It covers the three call types…
Last verified 2026-07-22 · Primary source
2026-05-28
Claude Opus 4.8 lands — and it fixes what 4.7 broke
Anthropic shipped Claude Opus 4.8 on May 28, six weeks after the divisive 4.7 release. Generally available on the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and…
Last verified 2026-07-22 · Primary source
2026-05-21
Cursor's productivity study: 39% more PRs at the same revert rate
Cursor published the data behind the productivity claim — a study Suproteem Sarkar (University of Chicago, finance and applied AI) ran across tens of thousands of Cursor users.…
Last verified 2026-07-22 · Primary source
2026-05-19
Karpathy joins Anthropic's pre-training team
Andrej Karpathy announced on May 19 that he had joined Anthropic. The short bio for anyone who's lost track: OpenAI cofounder, then Tesla's director of AI from 2017 leading…
Last verified 2026-07-22 · Primary source
2026-05-19
Google's Gemini Omni: one model, any modality in, any modality out
Google unveiled Gemini Omni at I/O 2026 — a model family intended to create across modalities. The first variant, Gemini Omni Flash, accepts combinations of text, images, audio,…
Last verified 2026-07-22 · Primary source
2026-05-19
Gemini 3.5 Flash beats 3.1 Pro on the benchmarks that matter
The other I/O headline: Gemini 3.5 Flash — Google's new Flash-tier model — outperformed Gemini 3.1 Pro on three coding and agentic benchmarks Google highlighted: Terminal-Bench…
Last verified 2026-07-22 · Primary source
2026-05-18
Composer 2.5 makes the in-house model competitive
Cursor shipped Composer 2.5 on May 18 — the same Moonshot Kimi K2.5 base as Composer 2, retrained with 25× more synthetic tasks and a new technique they call Targeted RL with…
Last verified 2026-07-22 · Primary source
2026-05-14
NVIDIA's SANA-WM puts minute-scale world models on a single GPU
NVIDIA Labs dropped SANA-WM on May 14 — a 2.6B-parameter open-source world model that turns one image plus a 6-DoF camera trajectory into 60 seconds of controllable 720p video,…
Last verified 2026-07-22 · Primary source
2026-04-30
Cursor Security Review goes managed
The DIY security agents Cursor open-sourced earlier just turned into a product. Cursor Security Review is in beta on Teams and Enterprise plans, with two always-on agents you…
Last verified 2026-07-22 · Primary source
2026-04-30
Vercel opens the cloud agent stack
Cursor's SDK gives you agent infrastructure from an IDE company. Vercel's answer is different: an open-source reference implementation you can actually read. Open Agents is…
Last verified 2026-07-22 · Primary source
2026-04-29
Cursor SDK opens the harness up
Cursor shipped a TypeScript SDK on April 29 — the same runtime, harness, and models that power the desktop app, CLI, and web client, now available programmatically via npm…
Last verified 2026-07-22 · Primary source
2026-04-28
Cursor ships agentic security review
Cursor released four automation templates based on security agents it runs internally, including Agentic Security Review. That template runs on pull requests, posts findings as…
Last verified 2026-07-22 · Primary source
2026-04-23
GPT-5.5 lands a week after Opus 4.7 — and the vibe flips
OpenAI announced GPT-5.5 on April 23, seven weeks after 5.4 and seven days after Anthropic's Opus 4.7, with API availability following on April 24. It's rolling out on ChatGPT…
Last verified 2026-07-22 · Primary source
2026-04-17
The packaging pattern works for design too
The same pattern that works for code conventions works for design. ux-ui-agent-skills packages DTCG design tokens, Atomic Design component specs, WCAG 2.2 checklists, Nielsen…
Last verified 2026-07-22 · Primary source
2026-04-17
Claude Design joins Anthropic Labs
Claude Design shipped April 17 as a research preview — Anthropic's first dedicated visual creation tool, running on Opus 4.7, with direct handoff to Claude Code for development.…
Last verified 2026-07-22 · Primary source
2026-04-17
Hooks move agents from advice to automation
Skills tell an agent what to do. Hooks make certain things happen regardless of what the agent decides. They fire on lifecycle events — before a tool executes, after it finishes,…
Last verified 2026-07-22 · Primary source
2026-04-17
Harness design is now part of the craft
A harness is everything around the model: prompts, tools, orchestration, context management, hooks. If you want the from-scratch primer on what a harness is and why it matters,…
Last verified 2026-07-22 · Primary source
2026-04-17
Subagents: the Cursor model is worth studying
Cursor's subagents go further than the AGENTS.md pattern. Each gets its own context window and model config, runs foreground or background, and three built-ins (Explore, Bash,…
Last verified 2026-07-22 · Primary source
2026-04-17
Next.js MCP is becoming practical
Next.js 16 ships with a built-in MCP endpoint at / next/mcp . Add next-devtools-mcp to .mcp.json and your agent gets live access to build errors, runtime errors, routes, page…
Last verified 2026-07-22 · Primary source
2026-04-17
Browser control is getting lighter-weight
The fastest way for an agent to use a browser is to let it write code. dev-browser runs Playwright-style scripts in a sandboxed QuickJS WASM environment — install it globally,…
Last verified 2026-07-22 · Primary source
2026-04-17
Infra is becoming part of the product
Cursor's self-hosted cloud agents are now generally available. A worker process connects outbound via HTTPS — no inbound ports, no firewall changes. Cursor handles inference and…
Last verified 2026-07-22 · Primary source
2026-04-17
Cloud agents run on your hardware now
Cursor's My Machines takes self-hosted agents from an enterprise feature to an individual one. Instead of running in Cursor's managed VMs, your agent executes on hardware you…
Last verified 2026-07-22 · Primary source
2026-04-17
Harnesses are becoming shareable infrastructure
everything-claude-code is a useful example of where this is heading: 67 specialized subagents, hooks for memory persistence, verification loops, continuous learning, and security…
Last verified 2026-07-22 · Primary source
2026-04-17
Engineering practices are becoming installable
Addy Osmani packaged Google's engineering culture into agent-skills: 24 skills across a 6-phase lifecycle, with eight slash commands ( /spec , /plan , /build , /test , /review ,…
Last verified 2026-07-22 · Primary source
2026-04-17
Design systems are going agent-readable
Google Stitch introduced DESIGN.md — a plain-text design system document that agents read to generate consistent UI. No Figma plugins, no design token APIs. Just a markdown file…
Last verified 2026-07-22 · Primary source
2026-04-17
Knowledge bases are replacing notebooks
Andrej Karpathy shared a pattern worth paying attention to: instead of using LLMs to write code, use them to build personal knowledge bases. The structure is simple. Raw…
Last verified 2026-07-22 · Primary source
2026-04-15
Canvases make agent output interactive
Cursor shipped canvases — interactive visual surfaces that agents create inline. Instead of reading a text-based summary of your data, the agent generates a custom dashboard,…
Last verified 2026-07-22 · Primary source
2026-04-07
Mythos Preview finds zero-days at scale
On April 7, Anthropic previewed Claude Mythos — an unreleased research model that's dramatically better at exploiting software than anything shipped before. On a Firefox…
Last verified 2026-07-22 · Primary source
2026-04-07
Glasswing: defenders get the model first
Mythos doesn't ship alone. Project Glasswing is the coordinated deployment — a partnership with 12 founding organisations (AWS, Apple, Google, Microsoft, NVIDIA, Linux…
Last verified 2026-07-22 · Primary source
2026-04-02
Open models keep closing the gap
Gemma 4 launched on April 2 with four variants; Google subsequently added a 12B model. The current family has five sizes: E2B, E4B, 12B, 26B A4B, and 31B. All use Apache 2.0 — a…
Last verified 2026-07-22 · Primary source
2026-04-02
Cursor 3: parallel agents and worktrees
Cursor 3 shipped several changes that address the same problem — waiting. The Agents Window and Agent Tabs let you run multiple agent sessions in parallel rather than treating…
Last verified 2026-07-22 · Primary source
2026-04-02
Visual annotation beats text descriptions
Cursor 3 shipped Design Mode. Instead of typing "change the third button in the second card on the settings page," you click on the element. Design Mode opens a browser panel…
Last verified 2026-07-22 · Primary source