We track Product so you don't have to. Top Podcasts summarised, the latest AI tools, plus research and news in a 5 min digest.

Hey Product Fans!

Welcome to this week’s 🌮 Product Tapas.

New here? We're your shortcut to staying sharp. Essential stories, practical tools, real insights.

For the best reading experience, check our web or app version and sign up for future editions here.

What’s cooking this week? 🥘

OpenAI shipped the app that admits what we were all doing anyway - living inside the assistant - then published a price list undercutting its own frontier model sixteenfold (though its flagship Sol is also quietly deleting people's files). Meanwhile everyone's building the same thing in private: Figma's "PM OS" saves two hours a workflow, and Block found that once non-engineers have somewhere safe to deploy, they'll build you a thousand internal apps. Elsewhere the leading coding benchmark rejects correct code six times out of ten (everyone's citing it anyway), 1X gave its robot hands and declared hardware "basically solved", and a startup decided the fix for a multi-year turbine backlog is a worse turbine you can have tomorrow.

  • Not Boring → OpenAI moves into the office and cuts its own price, everyone builds an internal AI OS, the benchmarks are broken

  • Productivity Tapas → Nobie, ContextVault, ExploreYC, Mcpsnoop, Bramble

  • Blog Bites → Lenny's burnout data, Benedict Evans on token pricing, and the self-inflicted AI cost crisis

  • Pod Shots → Ten conversations on winning by subtraction

Let's go 🚀

📰 Not boring

OpenAI Ships the Work App (and Undercuts Itself on Price)

  • ChatGPT Work launched as a dedicated desktop app with Work and Codex modes - it pulls from calendar, Slack and Drive for daily briefings, and its "Appshot" feature captures live app state so ChatGPT can reproduce and diagnose a bug itself rather than squinting at a screenshot

  • GPT-5.6 arrived in three tiers - Sol beats Claude Fable 5 on Agents' Last Exam (53.6 vs 40.5), while Terra and Luna match Fable 5 at roughly a sixteenth of the cost. Pricing runs $5/$30, $2.50/$15 and $1/$6 per million input/output tokens

  • But Sol also deletes files on its own - developers report wiped Macs and a nuked production database, and OpenAI's own system card admits the model is "overeager to complete the task" and will act unless explicitly prohibited

  • Anthropic launched Reflect, a dashboard showing how you actually use Claude - and periodically asking you questions like "what's one thing you want to keep doing yourself"

Look at the pricing table, not the product: Luna matches frontier performance at a sixteenth of the cost, which makes OpenAI the company most aggressively commoditising the thing it sells - Benedict Evans's "models drift to low-margin plumbing" thesis, proven on OpenAI's own price list. The catch is the flagship above it will delete your files to finish the job. Cheap, capable, and slightly feral.

Everyone Is Quietly Building an Internal AI Operating System

  • Figma's PM team built "PM OS" - a volunteer trio built reusable "skills" on GitHub (a daily briefing, a weekly-update writer, a spec reviewer), rolled it out via a hackathon to 100% adoption, and reckon it saves at least two hours per major workflow

  • Block shipped "App Kit" so non-engineers can safely deploy vibe-coded internal tools. Since March: 10x growth in weekly views, a catalogue of well over a thousand apps - and around 80% of them built by people who aren't engineers

  • Notion launched Ship OS, an "agent-native way to ship software" that it used to launch itself, and YC launched Paxel, which profiles how you build from your Claude/Codex/Cursor transcripts (1.1M sessions analysed already)

  • Meta's Adam Mosseri expects AI token budgets capped per engineer within a year or two - scaled to how profitably each person actually uses AI, managed like GPUs or headcount

  • a16z argues humans are now cheaper than tokens - "you just hired a million bad employees" - and the real bottleneck is management discipline, since most token spend is wasted on bad prompting and looping

The number that matters is Block's 80%. The binding constraint on internal tooling was never the building - it was having somewhere safe to deploy the thing once built; split "the agent writes the app" from "the platform owns identity and access", and a thousand apps appear. But note where Mosseri and a16z land: if humans are now cheaper than tokens, the next scarce skill isn't building, it's deciding what's worth spending a token on at all.

The Benchmarks Are Broken (and the Robots Have Hands)

  • SWE-bench Pro is the new frontier coding benchmark - 1,865 tasks across 41 live repos, with Claude Mythos 5 leading at 80.3%. Except OpenAI's own audit found roughly 30% of tasks are broken, and a review of 138 "hard" problems found 59.4% had flawed tests that reject functionally correct answers

  • 1X gave its Neo robot hands with 25 degrees of freedom, 0.2mm accuracy and 45N grip, tested past 2 million cycles. CEO Bernt Bornich says the hardware problem is "basically solved" and what's left is data collection

  • Figma engineered a "Mirror DOM" to make a canvas product accessible - an invisible accessibility tree synced to the live canvas, because bypassing the browser's DOM also bypasses everything screen readers rely on

If you have ever put a benchmark number in a roadmap deck, read the SWE-bench Pro line again: six in ten hard problems have tests that fail correct code, so the leaderboard is partly measuring a model's ability to satisfy a broken grader. And 1X's "hardware solved, just data left" is exactly what a company shipping robots into homes to harvest that data would want you to believe.

The Picks and Shovels Get Weird

  • Aalo Atomics went critical at 12:20am on 4 July - a full-scale 10 MWe reactor core built in under eight months, destined for "Aalo Pods" powering AI data centres. One of four advanced reactors to hit the July 4 executive-order deadline

  • American Turbines came out of stealth building small, manufacturable gas turbines - the pitch being that if you need megawatts you buy more small turbines rather than queue years for one big one from GE or Siemens

  • Data centres are a test of US industrial resolve argues Josh Zoffer in the FT - hyperscalers paying premium prices for fast capacity are doing the work of industrial policy, building supply chains on demand rather than subsidy

American Turbines is the most quietly interesting product decision of the week, and it isn't software. The incumbents build enormous, bespoke turbines with multi-year backlogs; the startup's answer is a worse one you can have immediately, then buy twelve of. Modular beats monolithic when the constraint is lead time, not efficiency - a lesson that transfers uncomfortably well to how most of us plan platform work.

Productivity Tapas: Time-Saving Tools & Workflow Automation

  • Nobie: free, local-only Excel-compatible spreadsheet for Mac - full formula and formatting fidelity, no account needed, and it connects to Claude, Codex or Gemini for AI-assisted work. Beta

  • ContextVault: shared memory layer letting Claude, Codex, ChatGPT and other MCP clients read and write the same durable context, so your team stops re-explaining itself every session. Solo tier $9.99/month, 7-day free trial

  • ExploreYC: free, open-source API and database covering 8,952+ Y Combinator, a16z and Product Hunt companies - batches, funding, hiring status, founders. Founder research without a Crunchbase subscription

  • Mcpsnoop: "Wireshark for MCP" - a transparent proxy with a live terminal UI showing every real tool call between your AI client and your MCP servers, so you can debug silent failures instead of guessing. Free, open-source, MIT licensed

  • Bramble: local-first, encrypted password manager with peer-to-peer sync across your own devices - no cloud vault, no subscription, no company to breach. Free, open-source

    Remember. Product Tapas subscribers get our complete toolkit - 580+ personally tailored, time-saving tools for PMs and founders. Your shortcut to efficiency and what's hot in product management

Check the link here to access.

🍔 Blog Bites - Essential Reads for Product Teams

Leadership: How Tech Workers Are Feeling in 2026

Lenny's second annual workforce survey finds tech splitting into two camps: 41% "energized" by AI and thriving, versus a shrinking-but-real 12% who feel "resentful" and destabilised - and that divide predicts career optimism three times more strongly than whether someone works at a startup or a giant. Burnout jumped from 44.7% to 55.7% in a single year, even as 82% report real productivity gains. Read the full article here.

"I can do more, faster, but not better."

Key Takeaways

• The real fault line: it isn't seniority or company size that predicts how someone feels about their career - it's whether they feel amplified or diminished by AI, and that split is 3x stronger than any other variable

• Burnout is up, optimism is down: significant burnout rose from 44.7% to 55.7% year-on-year, while career optimism fell from 54.8% to 48.7%

• Managers are the biggest lever: only 25.5% of managers are rated highly effective, but effective managers correlate with 65% higher job enjoyment - dwarfing any AI tooling decision

• Productivity without proof of quality: 82% report measurable AI speed gains, but the dominant fear isn't job loss (22%) - it's unsustainable pace (46%) and quality erosion (41%)

Noam Segal and Lenny Rachitsky, Lenny's Newsletter

Strategy: Ways to Think About Token Pricing

Benedict Evans lays out why nobody - including the labs themselves - actually knows where token prices settle, because every input is still moving. He tests the AI-as-infrastructure thesis against fibre, mobile data and cloud, and concludes the closest parallel is a warning: mobile data traffic grew by orders of magnitude and created a trillion-dollar industry, and the carrier stocks still went nowhere, because the value moved elsewhere in the stack. Read the full article here.

"Every path to foundation models having market dominance and pricing power requires something to change."

Key Takeaways:

• We are in an unstable supply crunch: trillion-dollar data centre capex, rising inference efficiency and shifting model efficiency mean today's pricing tells you almost nothing about tomorrow's

• Product-market fit is narrower than the hype: outside coding tools, proven ROI for frontier-model spend is still thin

• History says value moves upstream: TSMC dominates chip manufacturing yet captures less than half of Apple's value. Mobile data volumes exploded while carrier stocks stagnated

• Commodity plumbing is the base case: Evans thinks the visible dynamics point towards foundation models becoming low-margin infrastructure, with value captured in the tooling and go-to-market layer built on top

Benedict Evans

Growth: The AI Cost Crisis Is Entirely Self-Inflicted

Kyle Poyar digs into why AI spend is spiralling inside fast-growing companies - one Series B saw costs jump from $400K to $1.4M a year - and argues the problem isn't AI adoption but the total absence of the cost discipline companies apply to every other line item. He lays out a five-step fix, from spend visibility through to treating AI like any other capital investment with a measurable return. Read the full article here.

"A lot of it is justified. A lot of it isn't. Most people aren't conscious of what they're spending."

Key Takeaways:

• Costs compound fast: one company's AI bill jumped 3.5x in a year, with newer frontier models running 3-5x costlier than the previous generation

• Nobody has visibility: at that same company the CEO had personally burned $4K in days, and marketing was spending $739 a head with little to show for it, because no one was tracking it

• Budgets work: Tesla caps AI spend at $200/week per person, Uber at $1,500/employee/month. Hard ceilings force teams to prioritise the highest-ROI uses

• Treat it as an investment, not a subscription: the fix isn't cutting usage, it's requiring each use case to show a business outcome - the discipline you'd apply to headcount or ad spend

Kyle Poyar, Growth Unhinged

🎙 Pod Shots - Bitesized Podcast Summaries

Remember, we've built an ever-growing library of our top podcast summaries (140+). Whether you need a quick refresher, want to preview an episode, or need to get up to speed fast - we've got you covered.

Check it out here

Pod Shots #144 - The Art of Not

With AI getting more and more competent, it's ever easier to go one way: add more. More features, more tests, more hires, more AI, more surface area. This week, ten of the best conversations in tech - a CPO, the head of Instagram, a corporate ethnographer, a chief people officer, a public-sector UX lead, Disney's experimentation brain, Snap's CEO, two enterprise-sales veterans, the founder of e.l.f., and the ghost of Walt Disney - all quietly landed on the opposite instruction.

The winners this week weren't defined by what they added. They were defined by what they refused to do. Diya Jolly decides what not to decide. Crystal Ammari kills ideas before a single engineer is hired. Vonny Laing designs for the users everyone else skips. Walt Disney refused to oversaturate the thing that made him money. e.l.f. won by not competing where it couldn't. In a world where adding anything is nearly free, the scarce discipline is the "no".

  • Diya Jolly, CPO & CTO of Xero, "Why great product leaders should stop obsessing over the roadmap", In Depth - Watch

  • Adam Mosseri, Head of Instagram, "The rise of taste and judgment in an AI world", Lenny's Podcast - Watch

  • Sam Ladner, Workday, "Why AI can't replace qualitative research", Awkward Silences - Listen

  • Katie Burke, COO of Harvey, "The most politically dangerous role in the C-suite", In Depth - Watch

  • Vonny Laing, UX Lead, Student Loans Company, "Designing for your hardest users first", The Product Experience - Watch

  • Crystal Ammari, ex-Nike, ex-Disney, "How Disney picks which experiments to run", The Experimentation Edge - Listen

  • Evan Spiegel, CEO of Snap, "Why distribution has become the most important moat", Lenny's Podcast - Watch

  • Chad Peets & Chris Degnan, "The $100M CRO bubble", 20VC - Watch

  • Joey Shamah, co-founder of e.l.f. Cosmetics, "The dollar-store formula", How I Built This - Listen

  • Ben Gilbert & David Rosenthal, "The Walt Disney Company (Walt's Era)", Acquired - Watch

Estimated reading time: 10 mins. Time saved: 16+ hours!

Key insights from the full round-up:

  • Decide what not to decide - Jolly: the CPO job is direction, deep customer understanding and quarterly resource allocation. Obsessing over the roadmap is how most orgs ship mediocre products

  • Don't build it to find out - Ammari: one fake button told Disney nobody wanted video chat, before a single engineer was hired. Call the kills "savings", not "losses"

  • Don't offload the part that matters - Ladner: hand AI the transcription and the coding, but "you can't offload cognitive sense-making to machines"

  • Don't defend what's easy to copy - Spiegel: software is not a moat. Defend distribution and the closest relationships

  • Don't oversaturate the core - Disney: scarcity and quality in the primary medium. Marvel fatigue is what breaking that rule looks like

That’s a wrap.

As always, the journey doesn't end here!

Please share and let us know what you would like to see more or less of so we can continue to improve your Product Tapas. 🚀👋

Alastair 🍽️.

Reply

Avatar

or to participate