July 1, 2026
After a 3-week export-control ban, Fable 5 went live globally on July 1 — the price of reinstatement: Anthropic agreed to give the US government pre-release access to its future frontier models.[1]Anthropic: Redeploying Fable 5 The ban was triggered June 12 when Amazon researchers discovered a jailbreak that enabled routine defensive cybersecurity work; Anthropic now blocks the technique in 99%+ of cases.[2]Tech Brew: Anthropic's top models are back Alongside the return, Anthropic and partners Amazon, Microsoft, and Google are proposing an industry-wide framework for scoring jailbreak severity on four axes: capability gain, breadth, ease of weaponization, and discoverability.[1]Anthropic: Redeploying Fable 5 Critics including former Facebook CSO Alex Stamos argued the ban was counterproductive — a Chinese model now matches Mythos on specific cybersecurity benchmarks at far lower cost.[2]Tech Brew: Anthropic's top models are back
The US government imposed export controls June 12 after Amazon researchers found a bypass technique. Testing ultimately showed the jailbreak didn’t expose unique Fable capabilities — many less capable models could replicate the same behavior.[1]Anthropic: Redeploying Fable 5 The reinstated model has the bypass blocked but also restricts some legitimate cybersecurity tasks, and some coding/debugging work falls back to Opus for now.
Anthropic’s agreement with the US government includes: pre-release access to future national-security-relevant models, rapid jailbreak information-sharing protocols with US agencies, a dedicated joint government-industry AI security research team, and a new HackerOne bug bounty program specifically for cyber jailbreak submissions.[1]Anthropic: Redeploying Fable 5 The Rundown AI notes the updated safety filter “might also flag legitimate coding work,” but Anthropic claims “the vast majority of coding work is unaffected.”[3]The Rundown AI: Fable returns worldwide
The ban arrived amid intensifying AI competition with China. A Chinese model now matches Mythos on specific cybersecurity benchmarks at significantly lower cost, lending credibility to arguments that restricting frontier US models pushes organizations toward less-regulated foreign alternatives.[2]Tech Brew The hosts of Last Week in AI ~00:00 note nobody’s position on this “is entirely logically consistent” — including David Sacks’s framing, which conflated Fable (with guardrails) and Mythos (without) in public statements.
Theo points out the ban’s “export or reexport” language applied to API-hosted services, meaning T3 Chat and similar products serving international users had to suspend Fable access entirely during the shutdown — not just users in targeted regions.[5]Theo t3.gg: FABLE IS BACK! The timing is sensitive for Anthropic’s profitability goals and anticipated IPO.
A sharp hot take: Anthropic may be the most dangerous lab precisely because they’re genuinely committed to safety — a righteousness that drives them to take actions other profit-motivated labs would never dare.[6]Nerd Snipe: Is Anthropic actually the most dangerous AI lab?
~00:00 The argument: labs constrained by shareholder optics or simple profit-seeking have natural limits on radical action. A lab that genuinely believes it’s humanity’s salvation can rationalize moves that are harder to predict, challenge, or counterbalance. Noble intent doesn’t cap the scope of what gets justified in its name.
“They are genuinely concerned about safety. They are genuinely trying to make the safest possible model and do this right. And that nobility, that righteousness that they baked into themselves and their mission, drives them to do some pretty sketchy things.”
Claude Sonnet 5 launched June 30 at introductory pricing of $2/M input and $10/M output tokens — but benchmark tests reveal it uses 2–5x more tokens than competitors, making its real cost per task higher than Opus 4.8 or GPT-5.5.[8]Theo t3.gg: Breaking Down Sonnet 5 AICodeKing benchmarked it at 55.71% vs GLM 5.2’s 81.43% on his internal benchmark, costing $9.40 per run vs $4.88.[9]AICodeKing: Sonnet 5 Fully Tested The bright spot: it spontaneously orchestrates sub-agents — a behavior previously exclusive to Fable 5 — making it a compelling cheap worker in Fable-orchestrated multi-agent systems.[5]Theo t3.gg: FABLE IS BACK!
Anthropic lists introductory pricing at $2/M input and $10/M output through August 31, 2026 (then rising to $3/$15).[7]Anthropic: Introducing Claude Sonnet 5 But Theo clocked it using nearly 2x the tokens of Opus and ~5x those of GPT-5.5 on a standardized Xi benchmark test (69k tokens vs GPT-5.5’s much lower count). Artificial Analysis reports Sonnet 5 costs $2.29 per average agentic task — comparable to Fable 5 and higher than the previous Sonnet.[10]Artificial Analysis: Sonnet 5 agentic cost Theo’s read: “It honestly feels like Anthropic bumped up the tier for everything — what we used to use Haiku for, we now use Sonnet. What we used to use Sonnet for, we now use Opus.”[8]Theo t3.gg
~13:30 Theo’s core finding: Sonnet 5 spontaneously breaks work into sub-agents and stays on task across them — a behavior that was previously Fable 5’s distinctive capability. This earned it the “5” branding. His fish-game rebuild test showed Opus finishing in 27 minutes vs. Sonnet taking 2+ hours with a buggier result, so it’s not a replacement, but as a cheap worker under Fable orchestration it makes sense.
~21:30 Theo notes a dual-use request success rate drop from 97% (Sonnet 4.6) to under 92%, and a viewer-reported bug where raw thinking traces leaked on Claude.ai — one trace contained 21 instances of “let me” from the model’s internal reasoning.
Multiple coding tasks failed in the review: a 3JS contact lens model wouldn’t load even after several fix attempts, a bow-and-arrow game remained broken, and SVG generation was notably weaker than GLM 5.2. A weird agentic behavior: Sonnet 5 tries to work in the root /tmp directory instead of the current working directory, causing constant permission prompts.[9]AICodeKing: Sonnet 5 Fully Tested
Anthropic benchmarks show improvements on Terminal Bench, HLE, OS World Verified, and GDP Eval. The model is available as the default for Free and Pro users, with introductory pricing through August 31. Safety profile: lower hallucination rates, better prompt injection resistance, deliberately weaker than Opus at developing software exploits.[7]Anthropic: Introducing Claude Sonnet 5
Nate Herk digs through Anthropic’s internal system prompts and surfaces 8 prompting techniques that engineers use on Fable 5 itself — including a Fable-specific gotcha: asking the model to “explain its reasoning” silently reroutes your request to Opus 4.8.[11]Nate Herk: How Anthropic Engineers Actually Prompt Fable 5
Fable 5 is priced at $10/M input and $50/M output tokens (2x Opus), with promotional access via Claude plans ending July 7.
Anthropic launched Claude Science beta on June 30 — a macOS/Linux application for researchers that integrates 60+ curated scientific tools, databases like UniProt and PDB, NVIDIA BioNeMo models, and auto job submission to HPC clusters and Modal GPUs.[12]Anthropic: Claude Science AI Workbench
Claude Science consolidates the fragmented toolchain of scientific research: analyze literature, execute multi-step experiments, produce auditable artifacts, and iterate on figures and manuscripts in one environment. A reviewer agent validates citations and calculations. Session forking lets you compare analytical approaches side-by-side.
Pre-configured tool coverage spans genomics, single-cell analysis, proteomics, structural biology, and cheminformatics. Database connections include UniProt, PDB, Ensembl, Reactome, ClinVar, ChEMBL, and GEO. Domain-specific models in the toolkit: NVIDIA BioNeMo’s Evo 2, Boltz-2, and OpenFold3.
Computing scales from local laptops through HPC clusters via SSH and Modal for on-demand GPU. Available to Claude Pro, Max, Team, and Enterprise users.
GLM 5.2 from Z.ai is a 750B-parameter open-weight model that nearly matches frontier closed models — and Z.ai’s lead scientist claims they’ll have a Fable-level open-weight model before 2027, roughly 6 months away.[13]Two Minute Papers: AI Just Entered A New Era
The host flags that Anthropic’s promise of model honesty conflicts with Fable’s practice of silently routing queries to less capable models: “Anthropic promised us that Claude would be honest and then introduced Fable, which depending on your question, could pass it to a different, less capable model without telling you about it. I do not consider that to be honest.”
“Not your weights, not your model.”
The lead scientist’s public claim: a Fable-level open-weight system before 2027. The host takes it seriously given the leap from GLM 5.1 to 5.2 happened in less than 3 months. Hardware needed to run it today: tens of thousands of dollars. But the community is already distilling it to smaller sizes.
Nate B Jones argues the Fable ban proved the risk of renting AI memory from providers: when your model goes offline, your context goes with it. His response: OpenBrain, a personal AI memory framework where ~80% of the setup can now be built by talking to Claude, with the remaining 20% (secrets, permissions, approvals) staying human-controlled.[14]Nate B Jones: I Built My Own AI Memory by Talking to Claude
The Fable ban and ChatGPT 5.6 lockout exposed a critical vulnerability: users who built workflows around provider-hosted context lost everything when the model went offline. Nate also references the Nikita/Lemonade insurance incident as evidence that handing sensitive memory to a third party creates real-world risk.
Three components:
As of June 2026, roughly 80% of the OpenBrain setup can be completed by talking to Claude or Codex — a stark contrast to February when users needed to manually configure SQL databases and CLI tools. The remaining 20% (secrets, permissions, approval gates) stays explicitly human-controlled.
“Rent intelligence, own memory.”
Models are commodities you swap; memory and skills are personal assets you own. The goal is model-agnosticity — when Fable goes down, you switch to GPT or GLM without losing your context.
Every’s Head of Consulting details how Codex enabled a non-technical operator to build a custom email triage app (trained on 150 sent emails), a family care OS in 13 hours, and a CRM enrichment system that did weeks of work overnight — while settling the build-vs-buy question: she replaced her vibe-coded CRM with Attio.[15]Every: How Every's Head of Consulting Uses Codex Every Day
Every’s internal AI agent “Claudia” reads emails, ingests meeting notes, tracks inbound leads, and maintains the CRM. It started as a Wizard-of-Oz prototype but has matured as model capabilities improved. The key insight: Claudia excels at executing standard operating procedures but still requires human direction for judgment, taste, and client-facing relationships. Every ended up hiring a human operations person alongside Claudia rather than instead of one.
“AI is really good at executing against a standard operating procedure. The question of taste and reaching for excellence still requires direction and managerial support.”
Despite being able to build a custom CRM (and actually doing so), Every switched to Attio because maintaining homemade tools accumulated data quality debt. The framing: “Software is like bones (deterministic structure) and LLMs are like a brain and ligaments (flexible, adaptive). Real software is a compilation of thousands of logical rules you don’t anticipate upfront — companies like Attio spend their entire existence gathering those rules.” OpenAI’s solutions engineers similarly use Codex for rapid customer demos rather than building their own tooling from scratch.[16]OpenAI: Codex for Solutions Engineers
“In the era of AI, you can build anything. The question is should you build and maintain whatever you actually build.”
Advice for organizations going AI-first: start with existing systems, define which tasks must stay human, standardize one process at a time before scaling.
World’s Fair 2026 (7,000 attendees) opened with a keynote on the “Loopcraft” philosophy of software factories — followed by companies including Microsoft, OpenAI, Z.ai, MiniMax, and Factory.com showing off the next wave of agent harnesses, including a Google demo that ran 93 sub-agents for under $1,000.[17]AI Engineer: WF2026 Keynotes
Microsoft’s Pablo Castro introduced a three-category knowledge framework (intrinsic/extrinsic/learned) and demoed Foundry IQ including agentic retrieval and Agent Optimizer. Claude is now generally available in Azure AI Foundry.
OpenAI announced the GPT 5.6 model family achieving 750 tokens/sec on Cerebras hardware, and revealed a fully open Codex stack: API → harness → Apps Server → plugins.
Peter Steinberger from OpenClaw described managing a manager-agent instead of running 10 terminals directly. His three-shift bottleneck model: 2024 was about tokens, 2025 was about compute, 2026 is about attention — human attention becoming the scarce resource in agentic pipelines.
Zishan Lee presented GLM 4.2 benchmarks placing it between Opus 4.7 and 4.8, alongside the Zcode harness and open-weight rationale. Z.ai positions open weights as a geopolitical counterbalance.
Thomas Wolf and Olive presented MiniMax M3: 400B parameters, native multimodal from step one (not bolted on), 1M context window via MiniMax Sparse Attention, and 300M+ users already on the platform. The architecture was partially designed by an intern.
Theresa from Factory.com detailed their full software factory stack: LLM routing achieving 25%+ cost savings, “missions” (orchestrator/worker/validator sequences), a deferred context engine, and an agent readiness framework for production deployment.
Kevin How from Google’s Anti-Gravity team demo’d Anti-Gravity 2.0’s three 2026 primitives: dynamic sub-agents, sidecar triggers, and generative UI. The OS kernel demo used 93 coordinated sub-agents for under $1,000 total cost.
Sarah from Notion presented a model-agnostic playbook they call “Token Town,” arguing that companies should treat open-weight models as negotiation leverage against closed providers: “your supplier is your competitor.” Live multi-agent factory demo included.
A Grapile data analyst presented: agent-authored pull requests grew from under 1% to over 25% in one year, with quality metrics statistically indistinguishable from human-authored PRs.
Vibov presented BAML (Boundary AI Markup Language), an agent-first programming language with zero-cost execution tracing, inferred error types, cross-language FFI, and a “slop philosophy” that embraces LLM non-determinism rather than fighting it.
Kestra raised $25M on a simple bet: replace Python-heavy Airflow DAGs with YAML configs where tasks run in any language (Python, Node, Bash, SQL, containers) in a single pipeline. 2B workflows run in 2025 (20x YoY) with customers including Apple, JP Morgan, Toyota, and Bloomberg.[18]Better Stack: I Tried the Tool Trying to Kill Apache Airflow (Kestra)
Kestra’s core insight: a pipeline should be configuration, not code. The browser editor shows a live diagram of tasks lighting up as they execute, with timeline and per-step log views. The visual builder and code stay bidirectionally in sync — edit one and the other updates automatically. Triggers are built-in: schedule (cron), webhook, file-landing, or API call.
Vs. Airflow: Your pipeline becomes a readable YAML config that non-Python engineers can review and approve in PRs. Kestra’s engine is reported to parallelize better.
Vs. Zapier/Make: No SaaS, no per-task billing, self-hosted, built for real dev infrastructure.
Vs. plain Cron: Built-in retries, timeouts, dependency maps, and a full UI.
Caveats: Java/JVM backend needs ~4GB RAM minimum; YAML struggles with complex dynamic branching that Python handles better; enterprise features (SSO, RBAC, audit logs) sit behind the open-core paywall. Free tier gives you one shared login.
Kestra runs natively on Apple Silicon. Spin up locally with a single Docker run command. Growth figures (2B workflows, 20x YoY) are company-reported, not third-party audited.
A clean 2-minute explainer on prompt caching: by storing the KV cache state for static prompt prefixes (system prompt, tools, conversation history), models skip recomputing those tokens on every turn — making cache hits 10x cheaper than cache misses in tools like Claude Code.[19]Arjay McCandless: Prompt Caching
At inference time, every token in the prompt is converted to query, key, and value vectors. Without caching, the model recalculates all key-value vectors for the full context (system prompt + tools + history) every single turn. With prompt caching, the KV vectors for the static prefix are written to a cache and reloaded directly on subsequent turns — only the new tokens need fresh computation.
“If you’re using something like Claude Code, a cache hit is 10 times cheaper than a cache miss.”
Google’s June 2026 AI roundup headline: computer use is now in Gemini 3.5 Flash, letting custom agents see and act across desktop, mobile, and browser. Also: Nano Banana 2 Lite (fastest Gemini image model), Gemma 4 12B running locally on 16GB laptops, Android 17, and the world’s first AI arts museum.[20]Google Blog: June 2026 AI Updates
Midjourney — the image AI company with 40 employees and $200M revenue — is entering preventative healthcare with a proprietary ultrasound device that’s faster than MRI, significantly cheaper, and designed to scale to 1 billion scans per year.[21]Nate B Jones: Midjourney breaks into healthcare
Midjourney is leveraging its unusual financial position (highly profitable, self-funded, 40 people) to invest in a new domain. The device is a special kind of ultrasound that images the inside of the body — affordable, much faster than an MRI, and designed to be experienced like “a visit to the spa.” The goal: a billion scans per year within a few years, positioned as the biggest medical imaging breakthrough in 50 years.
June 2026’s GitHub trending is dominated by AI coding-agent infrastructure: plugins that curb over-engineering, cross-session memory, meta-harnesses running Claude Code + Codex + Cursor simultaneously, and an agent-first programming language from Baidu.[22]Github Awesome: GitHub Trending Monthly #8
The S&P 500 gained 14.9% in Q2 2026 — its best quarter since 2020. Semiconductor stocks achieved their best quarter ever (+92% for the Philly Semiconductor Index). Meanwhile Sherwood coins the “Terrific 10” — Mag 7 suppliers (Sandisk, Micron, Dell, Marvell, Applied Materials, etc.) who are outperforming their customers by riding the AI capex wave.[23]Sherwood / Snacks: The Terrific 10
The Terrific 10 — Sandisk, Seagate, Dell, Micron, Corning, Western Digital, Flex, Marvell, Applied Materials, Intel — supply components and services to the Mag 7. The thesis: “your massive capex are someone else’s massive sales beat.” As Apple, Microsoft, Amazon, Google, and Meta spend tens of billions on AI data centers, the money flows through to their suppliers.[23]Sherwood: Terrific 10
S&P 500: +14.9% in Q2 (best since 2020), +9.55% YTD. Dow: +8.85% YTD (best first half since 2021). Nasdaq: +12.8% YTD. Russell 2000: +20%+ YTD — best since 1991, largely driven by AI-related chip companies. Philadelphia Semiconductor Index: ~92% quarterly return — strongest ever.[24]Morning Brew: Stocks best quarter since 2020
Losers: Gold saw its worst quarter since 2013. Headwinds: anticipated Fed rate hikes, OpenAI’s delayed IPO, Nvidia’s slower growth suggesting the AI rally may be moderating.
Sherwood also notes developers at Nvidia, OpenAI, and GitHub are implementing a “caveman plugin” that makes AI tools communicate in simplified language, reportedly reducing Claude Code token consumption by ~75%.
A long-form interview with Kent Beck (inventor of TDD, XP, and co-creator of JUnit) on AI’s impact on software engineering. His hot take on Dario Amodei’s “coding is going away” claim: “that’s a statement by someone who doesn’t understand software engineering.”[25]Pragmatic Engineer: How Kent Beck shapes the software engineering industry
Beck’s central AI thesis: “We are accumulating code faster than we are accumulating trust.” Coding is fundamentally a trust-building and understanding-building activity — not just text generation. AI can generate the text, but humans still have to understand and be responsible for what gets shipped.
On TDD in the AI era: agents can self-validate, and the spec-before-generation pattern maps naturally onto TDD — write the test first, let the agent make it pass. He’s excited about using AI to pursue ideas shelved for 40 years, including a B+ tree in Go that outperforms Rust’s standard library.
On the industry playbook: “Nobody knows” is his explicit stance — “we are back in explorer mode after 20 years of extract mode.” TDD origin story: traced to a book about writing output tapes before programs, first implemented as SUnit in Smalltalk, JUnit co-written with Erich Gamma on a Vienna-to-Dulles flight.
“That’s a statement by someone who doesn’t understand software engineering.” — Kent Beck on Dario Amodei’s claim that coding will go away
Grant Sanderson (3Blue1Brown) tells Dwarkesh Patel that he initially thought AI would handle theorem-proving while humans focused on explanation — but now believes AI will excel at both, likely surpassing most mathematicians at the “explaining and distilling half.”[26]Dwarkesh Patel: AI That Discovers Math Will Also Explain It Better Than Us
Sanderson notes that the greatest mathematical thinkers — Einstein, Shannon, Feynman — were also remarkably lucid expositors. Their papers were readable, not impenetrable. He initially assumed AI would automate theorem-proving while human mathematicians shifted to the explanation/distillation work. But the same capability that discovers novel proofs likely produces clear explanations too. What remains for humans is the narrow edge that even those giants had: genuinely novel problem-solving ideas that nobody else thought of.
“I kind of suspect that actually they’ll also be like quite good at doing that and probably just like better than most humans are at like doing the explanation half and distilling half.”
Fireship’s rapid-fire history from ARPANET (1969) through TCP/IP, the World Wide Web, browser wars, dot-com crash, Google, Web 2.0, Facebook, and iPhone — ending with “two guys named Sam and Dario ingested the entire internet while tricking the rocks into thinking harder, and now they’re renting it back to us at a premium.”[27]Fireship: The weird history of the internet
~00:00 Key beats: Paul Baran’s packet switching (1960s), ARPANET’s first message (UCLA to Stanford, crashes after two letters — 1969), Ray Tomlinson inventing email and the @ symbol, Vint Cerf and Bob Kahn’s TCP/IP on Flag Day (Jan 1, 1983), Paul Mockapetris’s DNS, Tim Berners-Lee’s World Wide Web at CERN (boss’s feedback: “vague but exciting”), Marc Andreessen’s Mosaic/Netscape, Microsoft bundling IE to kill Netscape, AOL carpet-bombing CDs, 56k dialup, Napster giving computers AIDS, the dot-com bubble popping March 2000, Google PageRank, Ajax/Web 2.0, Facebook, the iPhone (2007), and now AI.
“Two guys named Sam and Dario ingested the entire internet while tricking the rocks into thinking harder. And now, they’re renting it back to us at a premium.”
A quick Pragmatic Engineer clip: the CAP theorem’s famous “pick any two of consistency, availability, and partition tolerance” framing is hand-wavy and incomplete — a view validated by Martin Kleppmann’s influential blog post, which was itself controversial.[28]Pragmatic Engineer: Martin Kleppmann also hated the CAP theorem
CAP theorem states distributed systems can guarantee only two of three properties: consistency (data stays in sync across nodes), availability (both servers respond to reads/writes), and partition tolerance (system behaves when nodes disconnect). The “two of three” framing is widely taught but the host argues it’s imprecise and the whole theorem is incomplete — a view independently expressed by Martin Kleppmann (“Designing Data-Intensive Applications”) in a blog post that sparked debate. Smart commenters pushed back, arguing Kleppmann was technically right but nitpicking.