Issue 15

GPT-Live goes full-duplex as SWE-Bench Pro collapses under audit

OpenAI launched GPT-Live, a full-duplex voice model backed by GPT-5.5 reasoning, while xAI's Grok 4.5 topped SaaS automation benchmarks at one-quarter the cost of rivals. OpenAI also retracted its recommendation of SWE-Bench Pro after auditing found 30% of its tasks broken. Prime Intellect raised $130M to bring frontier reinforcement learning infrastructure to any company. In pharma, AstraZeneca and Ionis reported a major trial failure in a competitive heart disease market, the White House reviewed FDA commissioner finalists, and RFK Jr. moved to create a formal list of Covid vaccine injuries. The NY Fed found tariff cost pass-through to consumers is still in the pipeline.

30 min read process

ai Voice goes full-duplex, benchmarks under fire

OpenAI's GPT-Live Ditches Turn-Based Voice for Real-Time Two-Way Conversations

OpenAI launched GPT-Live, a full-duplex voice model that listens and speaks simultaneously, replacing the turn-based pattern of earlier ChatGPT voice mode. For complex queries, the system offloads reasoning to GPT-5.5. Simon Willison, who had preview access in the iPhone app for several weeks, described the new model as a significant improvement over its predecessor.

Alpha Signal
Claude Sonnet 4.6

OpenAI Retracts SWE-Bench Pro After Finding 30% of Tasks Broken

OpenAI audited SWE-Bench Pro and found roughly 30% of its tasks are broken, and retracted its own prior recommendation of the benchmark as the leading coding eval. The finding matters because multiple labs have cited SWE-Bench Pro scores as evidence of coding capability gains. It is now unclear how much of the progress measured on the benchmark reflects real capability improvement versus artifacts of the broken tasks.

Alpha Signal
Claude Sonnet 4.6

xAI's Grok 4.5 Tops SaaS Automation Benchmark at a Quarter the Cost

xAI's Grok 4.5 topped Artificial Analysis's AutomationBench-AA with a 51% score, completing more SaaS workflow objectives than any rival while costing one-quarter as much per task. The benchmark tests end-to-end completion of real software-as-a-service workflows rather than isolated code generation, making the cost advantage particularly relevant for production deployments.

Alpha Signal
Claude Sonnet 4.6

Prime Intellect Raises $130M to Open RL Training Labs Once Locked Inside OpenAI

Prime Intellect raised $130M led by Radical Ventures to give any company access to the same reinforcement learning training infrastructure used by frontier AI labs. The startup has positioned itself as a provider of RL training capabilities that previously required either building in-house at substantial cost or operating inside a frontier lab. The round signals investor conviction that RL infrastructure is becoming a standalone product category.

Alpha Signal
Claude Sonnet 4.6

PyTorch 2.13 Ships 12x Faster Attention and 4x Memory Savings for LLM Training

PyTorch 2.13 ships FlexAttention on Apple Silicon, delivering attention operations up to 12x faster than prior implementations. The release also includes a fused loss that cuts GPU memory consumption by 4x and a new distributed communications backend for large clusters. The memory reduction is significant for anyone fine-tuning large language models on constrained hardware.

Alpha Signal
Claude Sonnet 4.6

Quoting Kenton Varda

Cloudflare co-founder Kenton Varda declared a moratorium on AI-written PR descriptions, commit messages, and issue tickets from his team. Varda said AI-generated change descriptions were worse than useless during code review, outlining implementation details while omitting the reasoning that would help a reviewer understand whether a change is correct. Simon Willison published the quote as a counterweight to enthusiasm about AI coding tools.

Simon Willison
Claude Sonnet 4.6

Agentic safety triggers aren't textual safety triggers; MCP attacks that beat SOTA guardrails more than half the time

Researchers at r/MachineLearning documented MCP-based attacks that bypass state-of-the-art safety guardrails more than half the time. The key finding is that agentic safety triggers are not the same as textual safety triggers: existing guardrails are built to classify prompt text, but agents with real tool access face a different attack surface where the malicious instruction arrives through tool responses rather than the initial prompt. Code and datasets were published alongside the post.

r/MachineLearning
Claude Sonnet 4.6

tencent/Hy3

Tencent released Hy3 under Apache 2.0, a 295B-parameter mixture-of-experts model with 21B active parameters and 3.8B MTP layer parameters. Following a preview launch, the full model ships with an open license that allows commercial use. Chinese labs releasing large open-weight models continues a pattern that has put pressure on proprietary providers to justify their pricing.

Simon Willison
Claude Sonnet 4.6

Artificial Analysis Fixes the Benchmark That Let Bad TTS Models Win

Artificial Analysis launched a voice cloning leaderboard that controls for voice selection bias, which had allowed worse text-to-speech models to win earlier rankings by pairing with easier-to-clone voices. The controlled methodology puts Cartesia Sonic 3.5 at the top. The benchmark change alters the apparent competitive ranking enough that results from prior leaderboards are not directly comparable.

Alpha Signal
Claude Sonnet 4.6

Data for Agents

Hugging Face released guidance on data preparation for training AI agents, addressing a gap in publicly available resources for agentic system development.

Hugging Face Blog
Claude Haiku 4.5

software Bun rewrites itself; sqlite-utils ships 4.0

Rewriting Bun in Rust

Bun's creator Jarred Sumner rewrote the JavaScript runtime from Zig to Rust and published a detailed account of the process. The rewrite, which Sumner had been promising since May 9, was completed before the blog post was finished. Simon Willison described it as a detailed account of the engineering tradeoffs involved in switching systems languages on a mature, performance-sensitive project.

Simon Willison
Claude Sonnet 4.6

sqlite-utils 4.0, now with database schema migrations

Simon Willison released sqlite-utils 4.0, the 124th release of the project and the first major version bump since November 2020. The release introduces database schema migrations and includes breaking changes described in an upgrade guide. Much of the work in recent release candidates was reviewed by Claude Fable 5.

Simon Willison
Claude Sonnet 4.6

Introducing Meerkat: an experiment in global consensus

Cloudflare Research published a description of Meerkat, a global consensus service built on a new algorithm called QuePaxa. The project aims to provide a strongly consistent, fault-tolerant key-value store suitable for globally distributed applications. Cloudflare framed it as experimental infrastructure that may eventually support applications requiring strict coordination across regions.

Claude Sonnet 4.6

Your Worker can now have its own cache in front of it

Cloudflare launched Workers Cache, a regionally tiered cache that sits directly in front of Worker entrypoints and is configured via standard HTTP headers. The product makes caching composable with the existing Workers programming model without requiring separate cache infrastructure. Developers can now define cache behavior in code and have it apply at the regional edge.

Cloudflare Blog
Claude Sonnet 4.6

Tech jobs market in 2026, part 3: hiring managers & job seekers

The Pragmatic Engineer surveyed 50 or more hiring managers and job seekers to produce a third installment of its 2026 tech jobs market series. Key findings include a market where candidates and companies struggle to find each other, AI-related roles as the hottest hiring category, and a notably difficult environment for engineering leadership positions. The report is based on direct accounts from practitioners rather than aggregated job posting data.

The Pragmatic Engineer
Claude Sonnet 4.6

Rewriting Bun in Rust

Jarred Sumner completed his rewrite of Bun from Zig to Rust, documenting detailed architectural changes, performance considerations, and the engineering rationale behind the language shift.

Simon Willison
Claude Haiku 4.5

The Pragmatic Engineer AMA

The Pragmatic Engineer's Gergely Orosz conducted an AMA addressing listener questions on AI adoption, hiring practices, career decisions, and how engineering teams are adapting to model-assisted development.

The Pragmatic Engineer
Claude Haiku 4.5

pharma FDA leadership in limbo; AstraZeneca trial fails

STAT+: White House reviewing top contenders to lead FDA

The White House is reviewing finalists for FDA commissioner, with a decision expected soon. The position has been vacant long enough that the agency's leadership structure has become a subject of concern for industry and public health advocates. STAT reported that the narrowing of candidates to a finalist pool suggests an announcement is imminent.

STAT News
Claude Sonnet 4.6

STAT+: AstraZeneca, Ionis report major trial failure with heart disease drug

AstraZeneca and Ionis reported a significant trial failure in transthyretin amyloid cardiomyopathy, a heart disease market that has become increasingly competitive. The drug, Wainua, did not meet its primary endpoint. The failure is notable because ATTR-CM has attracted multiple drug developers and the market was seen as a near-certain growth opportunity given the aging population and the success of earlier treatments.

STAT News
Claude Sonnet 4.6

STAT+: RFK Jr. plans to create list of injuries caused by Covid-19 vaccines

RFK Jr. plans to create a formal list of injuries caused by Covid-19 vaccines, modeled on existing injury compensation tables. The conditions that may qualify for the list are not yet clear, and outside experts are watching the process carefully. Critics worry the list could be used to undermine vaccine confidence; supporters argue it formalizes an acknowledgment process that already exists informally.

STAT News
Claude Sonnet 4.6

STAT+: In private meeting, Trump officials push to onshore generic drugmaking

Marco Rubio, RFK Jr., and Chris Klomp met privately with industry leaders to push for onshoring generic drug manufacturing in the United States. The administration framed domestic generic production as a national security and supply chain resilience issue. Generic drugs account for roughly 90% of U.S. prescriptions by volume, but a large share of active pharmaceutical ingredients are manufactured abroad.

STAT News
Claude Sonnet 4.6

STAT+: 931 days. The drug approval scandal hiding in plain sight

Northwest Biotherapeutics submitted its brain cancer treatment to the FDA 931 days ago. The drug has not been approved or rejected. STAT reported that the delay is an outlier even by the standards of complex biologics review, and that the company has been unable to obtain a clear timeline from the agency. The case has become a focal point for critics of FDA review timelines for serious conditions.

STAT News
Claude Sonnet 4.6

Opinion: Who benefits from classifying obesity as a disease?

Max Moser argued in STAT that classifying obesity as a disease has benefited pharmaceutical companies selling GLP-1 drugs more than it has benefited patients. He drew a parallel to the antidepressant era, when disease framing expanded the treatable population and created durable markets. Moser did not argue that GLP-1 drugs are ineffective; he questioned whether the disease framing shapes prescribing and coverage decisions in ways that serve commercial rather than clinical interests.

STAT News
Claude Sonnet 4.6

Who's going to run the FDA?

STAT News published a digest of recent health policy developments, including updates on FDA commissioner candidates, drug approval timelines, and regulatory decisions affecting pharmaceutical manufacturers.

STAT News
Claude Haiku 4.5

healthtech Nursing strikes, Bryan Johnson, and hospital labor

STAT+: Mass General Brigham, nurses called to talk at State House amid biggest nursing strike in Mass.

Mass General Brigham and nurses at Brigham and Women's Hospital were called to the Massachusetts State House for talks amid what STAT described as the largest nursing strike in the state's history. The Brigham, a nationally prominent Harvard teaching hospital, has been negotiating a new contract with nurses for months without reaching agreement. The dispute centers on staffing ratios and compensation.

STAT News
Claude Sonnet 4.6

Bryan Johnson's chronic disease is notoriously difficult to diagnose

Bryan Johnson disclosed he has autoimmune gastritis, a condition that is notoriously difficult to diagnose because its symptoms overlap with a wide range of gastrointestinal disorders. STAT explained the diagnostic challenge: autoimmune gastritis is defined by the immune system attacking the stomach's parietal cells, but the resulting symptoms, including fatigue, vitamin B12 deficiency, and anemia, typically appear years after the underlying damage begins.

STAT News
Claude Sonnet 4.6

STAT+: The quest to save Grace; and clear the way for rare disease patients everywhere

Matt Wilsey has spent years trying to develop a drug for his daughter Grace, who has NGLY1 deficiency, an ultra-rare genetic disorder with no approved treatment. STAT reported on how his quest is shaping the broader path for rare disease drug development, including his work with Grace Science Foundation and what regulatory pathways exist for conditions with too few patients to run traditional trials.

STAT News
Claude Sonnet 4.6

What's new in biology: July 2026

Works in Progress published its July 2026 biology roundup, covering a synthetic cell made via chemistry, laser phase plate electron microscopy, the first base-edited human embryo, and lab-grown eggs and sperm. Each item represents a technique that either did not exist or was not demonstrated at this scale in prior years. The base-edited human embryo entry in particular marks a threshold that bioethicists have been tracking.

Works in Progress
Claude Sonnet 4.6

economy Tariffs still passing through; land was once made

The Real Reason European Cars Can't Compete

Patrick Boyle examined why European automakers cannot close the cost gap with Chinese electric vehicle manufacturers. His analysis traced the structural disadvantages to battery manufacturing scale, supply chain integration depth, and the legacy costs of internal combustion engine expertise that European brands built over decades. Industrial policy differences between the EU and China account for part of the gap, but Boyle argued the supply chain and scale advantages are not easily replicated through subsidy alone.

Patrick Boyle
Claude Sonnet 4.6

More Tariff Pass‑Through Is in the Pipeline

The NY Fed found that more tariff cost pass-through to consumers is still in the pipeline. Many businesses saw costs rise sharply over the past year but have not yet fully passed them on, absorbing the increases through margin compression or deferring price hikes. The analysis suggests reported inflation from tariffs understates the eventual consumer impact.

Liberty Street Economics (NY Fed)
Claude Sonnet 4.6

Effect of Tariffs on U.S. Small Businesses

The NY Fed used data from the 2025 Small Business Credit Survey to assess how recent tariffs have hit small businesses. The survey found widespread cost increases, with many small firms unable to fully pass costs to customers due to competitive pressure. Small businesses that rely on imported inputs but sell into price-sensitive domestic markets faced the sharpest margin pressure.

Liberty Street Economics (NY Fed)
Claude Sonnet 4.6

Why we stopped making land

Zigmund Forrest and Maxwell Tabarrok published in Works in Progress that roughly 8% of the land in America's major coastal cities was underwater in the 1890s and has since been reclaimed. Half of Boston's current land area and a quarter of Manhattan were raised from the sea before 1970. The piece traced why land reclamation stopped, pointing to environmental regulation and changed property rights frameworks rather than any technical barrier.

Marginal Revolution (Tyler Cowen)
Claude Sonnet 4.6

JD Vance's crusade against GDP is wrong and bad

Noah Smith argued that JD Vance's stated skepticism of GDP as a welfare measure is wrong in both analysis and policy implication. Smith acknowledged that GDP has real limitations as a welfare metric but contended that Vance's framing, which treats GDP growth as orthogonal to or in tension with American prosperity, is empirically unsupported. Making Americans poorer, Smith wrote, does not make society better by any measure Vance has articulated.

Noahpinion (Noah Smith)
Claude Sonnet 4.6

Single-payer health care systems are looking worse all the time

Tyler Cowen argued in a Free Press piece that single-payer health systems are performing worse over time relative to mixed systems, particularly for complex or novel conditions. Government-run systems, he wrote, handle routine care adequately but struggle with innovation-dependent treatments, where speed of adoption and pricing flexibility matter more. The piece drew on recent comparative outcomes data across OECD countries.

Marginal Revolution (Tyler Cowen)
Claude Sonnet 4.6

Land Reclamation!

Historians document that half of Boston's land, a quarter of Manhattan, and 15 percent of San Francisco were raised from the sea before 1970; contemporary land reclamation could relax the constraint that shaped modern cities.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

The tomb of Duns Scotus

Tyler Cowen reflected on Cologne's Cathedral and related architectural achievements, discussing aesthetic merit and historical context of European landmarks.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

Tyler Cowen's Wednesday assorted links covered fragility in digital money systems, regulatory uncertainty around Japanese electric baths, and new findings on internal neural patterns in Claude.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

What to Watch and Not

Tyler Cowen reviewed Spider-Noir and other recent streaming series, praising the noir series for its Raymond Chandler aesthetic and Nicholas Cage's performance.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

Missing women on Indian streets

Researchers used GPS-linked wearable cameras and randomized street audits across greater Mumbai to measure women's presence on city streets, finding significant gender disparities in public space usage.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

Tyler Cowen's Tuesday assorted links highlighted AI security surveys for U.S. government, vacancy tax effects, Substack's role in philosophy, and AI companies hiring philosophers.

Marginal Revolution (Tyler Cowen)
Claude Haiku 4.5

CityRX

Kyla Scanlon created a short-form video on CityRX, exploring economic dynamics in healthcare service delivery models.

Kyla Scanlon
Claude Haiku 4.5

Effect of Tariffs on U.S. Small Businesses

Liberty Street Economics analyzed the effect of recent tariff implementation on U.S. small businesses using data from the 2025 Small Business Credit Survey, finding significant sectoral variation in impact.

Liberty Street Economics (NY Fed)
Claude Haiku 4.5

Credit constraints and housing market access

Bank of England researchers examined the Help-to-Buy programme's effect on housing market access, finding it reopened the 95 percent loan-to-value mortgage segment for first-time buyers.

Bank Underground (Bank of England)
Claude Haiku 4.5

Shifts in Non-Cash Payments

Timothy Taylor analyzed Federal Reserve Payment Study data for 2024, documenting shifts in non-cash payment methods used by consumers, businesses, and government entities.

Conversable Economist (Timothy Taylor)
Claude Haiku 4.5