Lenny's Newsletter

▶ YouTubecompleted
Back to results

Lenny's Podcast · Jun 28, 2026

OpenAI Codex lead on the new shape of product work | Andrew Ambrosino

AI product developmentproduct processtaste and judgmentrole fluidityPM role evolutiondesign and AIdogfoodingproduct planningagentic workflowsprototypingfeature timingproduct discoveryOpenAI Codexzone defense PMimplementation abundance

Andrew Ambrosino, product and engineering lead for OpenAI's Codex app, argues that AI has inverted the traditional product process: implementation is now cheap and abundant, making taste, curation, and judgment the new scarce resource. The episode covers how product roles are evolving, why design lags behind code in AI capability, and how autonomous agent workflows are reshaping day-to-day product work.

Topics (13)

Process Inversion: Implementation is No Longer the Bottleneck

5:01

Because anyone can now stand up a feature by talking to a model, the constraint in product work has shifted from building to deciding what is worth building and keeping.

How it works

  • Previously, implementation was expensive, so teams front-loaded work: docs, research, prototypes were used to de-risk before a single line of code was written.
  • Now, implementation is abundant: 90 different teams inside OpenAI may independently prototype the same feature without coordination.
  • The bottleneck has moved to the curation layer: which of the 90 attempts is good, what should be folded together, how should it be framed?
  • The PM and design skills that matter most are now selecting, steering, and making sense of outputs — not unblocking implementation.
  • Why you should care

    If your product process is still organized around de-risking implementation, it is misaligned with the actual cost structure of your team. Your planning, approval, and resource allocation rituals were designed for a world that no longer exists.

    Think of it like...

    It is like moving from a world where printing a book costs thousands of dollars to one where printing is free — the job is no longer typesetting, it is editing.

    Taste as the New Core Product Skill

    11:18

    Taste is the ability to make high-quality decisions across aesthetics, systems thinking, and strategic direction when there is no shortage of things to choose from.

    How it works

  • Aesthetic taste: whether an animation or interaction pattern semantically fits what it is trying to convey.
  • Systems taste: understanding how a feature fits within the broader product architecture and theme.
  • Strategic taste: knowing what to work on at all, given unlimited capacity to build anything.
  • Medium taste: choosing whether a given problem needs a document, a prototype, or a direct experiment.
  • Why you should care

    This is the argument that traditional PM process skills (writing PRDs, running sprints) are necessary but no longer sufficient. The PMs who thrive are those who can stand in a room full of polished prototypes and make a confident, well-reasoned call on what is right.

    Think of it like...

    Taste is the editorial function in a newsroom flooded with contributed articles — the value is not in writing more pieces, it is in knowing which one to run.

    Primal Mark: The Anchoring Risk of Prototypes

    8:28

    Whatever medium you use to express an idea first shapes every decision that follows, so the choice of starting format is itself a high-stakes product decision.

    How it works

  • In art and design, the first mark on a canvas sets the reference point for everything after — teams respond to it rather than to the underlying idea.
  • A polished prototype can look production-ready but actually represent a very early-stage hypothesis, misleading stakeholders into treating it as a shipping decision.
  • Previously, the medium carried implicit stage information: a Figma file meant design was done, a working demo meant engineering had reviewed it.
  • Now that implementation is cheap, a fully functional prototype may have been built in hours and represent zero validated learning.
  • Why you should care

    Stakeholders — including executives — will over-anchor on whatever they see first. If you show a prototype before the problem is defined, you have effectively made a product decision by accident.

    Think of it like...

    It is like showing someone a rough sketch versus a finished rendering of a house — the rendering makes them react to details that may not even belong in the design yet.

    Medium Selection as a First-Class Product Decision

    7:52

    Because you can build anything quickly, you must now explicitly decide which format best serves the specific point you are trying to make — that decision is no longer implicit.

    How it works

  • Historically, medium correlated with stage: documents meant early exploration, prototypes meant design-done, code meant engineering-reviewed.
  • AI collapses that correlation: a working app can now be produced at the same stage that a sticky-note sketch used to be.
  • The right medium depends on the goal: a document for alignment around a vague area, a prototype to stress-test an interaction, a live experiment to measure behavior.
  • Engineers now tend to over-document; non-engineers tend to over-prototype — both are misfires driven by capability rather than appropriateness.
  • Why you should care

    If your team defaults to prototypes because they are easy, you are making a format choice that signals false maturity to every stakeholder who sees it. Explicit stage-setting becomes a communication responsibility that did not exist before.

    Think of it like...

    It is like choosing between a voice memo, a text message, and a formal email — the content might be the same but the format changes what the recipient does with it.

    Role as Average of Time Spent (Role Fluidity)

    23:28

    Your real role is the weighted average of your daily activities, not the label on your org chart — and that average is now expected to shift continuously.

    How it works

  • At OpenAI's Codex team, designers write code and contribute to product decisions; PMs write code and work on design questions.
  • Roles are defined by overlap and center of gravity, not by hard boundaries between functions.
  • The boundary of 'this is not your lane' is dissolving; what remains is depth and best-practice ownership within a discipline.
  • This does not mean roles disappear — it means the fence around a role matters less than the average contribution pattern.
  • Why you should care

    For a PM, this is both an opportunity and a threat: you can contribute more broadly, but you are also competing with engineers and designers who are expanding into your traditional territory. Your defensible value is judgment and depth, not role monopoly.

    Think of it like...

    It is like a basketball player whose position is described by where they spend most of their time on the court, not by a fixed label — a 'power forward' who mostly plays guard is effectively a guard.

    Zone Defense Product Management

    29:08

    Zone defense PM means each product person takes responsibility for a patch of the product landscape rather than for a specific team or process, filling gaps as they appear.

    How it works

  • In environments where everyone builds, the bottleneck is not execution but coherence — things get built without a shared direction.
  • PMs spread out to maximize coverage, asking 'where are the gaps?' rather than 'what is my feature?'
  • Two PMs working too closely is a signal of wasted coverage; the goal is company-wide product coherence with minimal overlap.
  • The PM hires for product-minded engineers so the coverage problem scales without requiring a PM for every output.
  • Why you should care

    This reframes what a PM's day looks like in an AI-forward org: less roadmap ownership, more pattern recognition across many parallel efforts. If you are used to owning a track end-to-end, zone defense requires a mental shift toward influence without direct ownership.

    Think of it like...

    It is like a basketball zone defense where each defender owns a region of the court, not a specific opponent — you go where the ball is, not where your man is.

    Hazard of Eliminating Product Roles

    1:12

    When companies abolish the product role, they do not just lose headcount — they lose the institutional knowledge of what has been tried, failed, and why.

    How it works

  • The appeal: if everyone can build, dedicated product people seem redundant.
  • The problem: product management as a discipline has decades of hard-won knowledge about what works — discovery methods, prioritization frameworks, stakeholder management — that is not obvious to engineers who have never practiced it.
  • Engineers who can vibe-code a feature often underestimate that knowing what to build and whether it is the right thing is a separate, learnable skill with real best practices.
  • Eliminating the role eliminates the accountability structure for those decisions.
  • Why you should care

    This is a direct risk for PMs in orgs moving fast on AI adoption. The counter-argument is not 'PMs are special' — it is 'product discipline encodes knowledge that is expensive to rediscover.' Make that knowledge visible and explicit.

    Think of it like...

    Getting rid of PMs because engineers can now write code faster is like getting rid of editors because journalists can now publish instantly — the constraint was never the printing press.

    Feature Timing Over Feature Shape (Model-Readiness Dependency)

    33:31

    In AI products, timing a feature release to model capability maturity is as important a product decision as the feature's design.

    How it works

  • The original Codex CLI was 'too AGI-pilled' — the product shape assumed model capability that did not yet exist, so it failed in the market.
  • Claude Code launched with a more constrained, interactive shape matched to actual model capability at that moment, and it worked.
  • The Codex app released in February 2025 would have failed if released in November 2024 — the only difference was a few months of model improvement.
  • Strategy: prototype features that are not yet model-ready, let them sit, and re-test each time there is a capability leap.
  • Why you should care

    Traditional product logic says a failed feature is a signal to kill it or pivot. In AI products, failure may be a timing signal, not a product-market fit signal. Your backlog of 'failed' AI features deserves a second look every model generation.

    Think of it like...

    It is like launching a streaming video service in 1998 — the product idea was right but the infrastructure was not ready; the same launch in 2010 is Netflix.

    Baby Product / Simplified Codebase for Design Exploration

    20:18

    A 'baby product' is a stripped-down replica of your production app that approximates all key interactions, purpose-built for fast, low-risk design exploration.

    How it works

  • The baby codebase mirrors the production app's interaction patterns but removes production complexity, making it fast to modify via AI coding tools.
  • Designers and PMs can explore 'what if the sidebar worked like this' without touching live infrastructure or waiting for engineering capacity.
  • It replaces the role that Figma interactive prototypes used to play, but with real running code.
  • Changes are not intended for production — the artifact is an exploration and comparison tool, not a shipping candidate.
  • Why you should care

    This is a concrete process primitive that replaces Figma-based prototyping for teams with AI coding capability. If your design process still routes through static prototypes for interaction exploration, this is a more testable and representative alternative.

    Think of it like...

    It is like a wind tunnel model of a car — physically accurate enough to test aerodynamics, but never intended to drive on a road.

    Why AI Is Weak at Design: Gradeability and Novelty

    12:53

    Design capability lags in AI models because creating a feedback loop that teaches a model what good design is requires human taste as part of the grading mechanism, which is harder to automate than checking if code compiles.

    How it works

  • Code has objective correctness signals: does it compile, does it pass tests, does it do what was specified? Design lacks an equivalent.
  • Labs historically invested in capabilities that accelerate AI research — correct code does that; good design does not directly.
  • Design also requires novelty: a model that always outputs the current best-practice design (e.g., a Linear-style website) would be producing the right answer for the wrong reason.
  • There is also a deep abstraction layer: visual design is entangled with code architecture — semantic relationships between components are not visible but must be preserved.
  • Why you should care

    AI-generated design will remain unreliable for brand differentiation and novel UX patterns for the foreseeable future. Design taste and originality are defensible human contributions — do not automate this prematurely.

    Think of it like...

    Training a model on design is like training a judge by giving them a rulebook with no cases — the judgment required is relational and contextual, not rule-following.

    Dogfooding Loop as Product Discovery Engine

    22:59

    Deliberately using your own product for real work, even when it slows you down, generates the most honest and actionable product feedback possible.

    How it works

  • The Codex team used the Codex app to build the Codex app — every friction point encountered was a real bug or gap to fix.
  • This creates a tight personal feedback loop: you hit a problem, you fix it, you can now do more, which reveals the next problem.
  • The team deliberately chose not to optimize their process in order to force themselves to improve the product — accepting short-term inefficiency for long-term product quality.
  • It also aligns the team's incentives: everyone wants the product to be better because they personally suffer when it is not.
  • Why you should care

    Structured research and analytics surface what users report; dogfooding surfaces what users experience but do not articulate. In fast-moving AI product development, the latency of traditional research cycles is too high — dogfooding compresses discovery to near-zero.

    Think of it like...

    It is like a chef who only eats at their own restaurant — they will find every bad dish faster than any critic.

    False Precision in Long-Horizon Planning

    32:10

    Any precision added to a roadmap beyond a short time horizon in an AI-driven product context is fabricated — it signals confidence that does not exist and costs real planning time.

    How it works

  • Short-term plans (weeks) can carry detail because the model landscape and product state are relatively stable.
  • Longer-term plans (months) should stay intentionally hazy — anything that looks specific is false precision.
  • The right long-horizon approach: list things you are interested in, prototype them, let the non-ready ones sit, and re-evaluate after each model capability leap.
  • Features that seemed viable may become wrong as model capabilities change; features that seemed impossible may suddenly work.
  • Why you should care

    This challenges the enterprise product planning ritual of quarterly and annual roadmaps. For AI product work in particular, a detailed 9-month plan is not a commitment — it is a liability that anchors teams to assumptions that will be wrong.

    Think of it like...

    Adding detail to a 9-month AI product plan is like drawing a precise weather map for next autumn — the framework is useful but the specifics are fiction.

    Do Not Get Married to Your Process — Get Married to Your Outcomes

    1:08:15

    The sustainable career move in an AI-accelerated environment is to hold your process loosely and your outcomes tightly — tools and steps are replaceable, the value you deliver is not.

    How it works

  • People who define their effectiveness by mastery of a specific tool (Figma auto-layout, TypeScript syntax) become vulnerable when AI makes that tool trivial.
  • The gatekeeping layer of many roles — the part that required deep tool expertise — is eroding fastest.
  • What survives: the ability to enter a problem space, learn what works, focus on it, and deliver outcomes.
  • This is not anti-specialization — it is pro-outcome: go deep on things that matter, stay loose on how you get there.
  • Why you should care

    For any PM with 10+ years of practice, the process you have refined is also the process most likely to be partially automated. The question is whether your identity is in the process or in the judgment that underlies it.

    Think of it like...

    It is the difference between a navigator who knows how to read stars and one who knows how to use GPS — when satellites go down, one of them is still useful.

    Why this matters to you

    Your product process is built for a cost structure that no longer exists.

    Every gate, approval step, and de-risking ritual in traditional product development was designed because implementation was expensive. AI has made implementation abundant. If your process has not changed, your team is paying a tax for a constraint that is gone.

    The PM role is not dying — but the defensible part of it is shifting fast.

    The parts of PM work that were gatekept by access to implementation (writing specs, coordinating design/dev handoffs) are dissolving. What remains — and what cannot easily be automated — is judgment: what to build, when, in what form, for whom. That is where PM value concentrates now.

    AI product failures may be timing failures, not product failures.

    If you have features in your graveyard that failed because the AI was not good enough, those are worth revisiting with every model generation. The product shape may have been right; the model readiness was not. This is a different backlog management mental model than traditional PM practice.

    Relevant transcript passages (5)
    4:06
    it's been kind of research ideation maybe there was some prototyping but it was you know even when we got past waterfall it was still kind of flavored of like the implementation is expensive and so you what you want to do is you want to derisk all implementation up front through documents through research through prototypes because prototypes and designs are cheaper was kind of the the assumption there uh and that's changed that's like totally changed

    Concise historical framing of why the old product process was structured the way it was — and the explicit claim that this assumption has now been invalidated. Directly relevant for any PM questioning whether their current process still makes sense.

    33:31
    I am very confident that the Codex app that we released in February, if that had been ready in November, it would have absolutely failed in the market and that that the only difference was the models between November and February, right?

    The clearest single data point in the episode on model-readiness as a product variable. Directly actionable: it reframes how to interpret AI feature failures in your own backlog.

    14:56
    there's sort of a an abstraction layer that is an interplay between the software design and the code it's being written. Like this thing over here in this corner should share x y and z in the codebase with this thing down here. Right? And that's a little bit different than saying the model needs to be a better designer

    Explains why AI design weakness is not just about visual aesthetics but about deep semantic relationships in codebases — relevant for anyone thinking about where AI-generated UI will and will not be reliable.

    31:19
    I generally look for like obviously command over the discipline but then the taste to say like hey you're going to have unlimited tokens and I don't like we can't just be doing slop like you need to be able to determine what's signal, what's noise, like in a world of just infinite content.

    Operationalizes what 'taste' means in hiring terms at an AI-native product team — useful signal for what to screen for when building or joining such a team.

    54:14
    the lesson in all of this was just that like the whole developer tool versus general knowledge work tool like there's a lot of nuance here that isn't just one or the other. And I think we really we believe really strongly in this and that there are certainly in the same way that we talk about the average of your role is like what your role is now. This is true on the product side too.

    Connects the role-fluidity concept to product positioning — the same 'average' framing that applies to people applies to products. Relevant for product strategy and positioning decisions.

    Key Insights (10)
    • Implementation is no longer the bottleneck in product development — curation, judgment, and taste are. Teams must reorganize around this new cost structure.
    • The medium you choose to communicate an idea (doc vs. prototype vs. experiment) is now an explicit product decision, not an implicit one — because medium no longer reliably signals stage or fidelity.
    • AI models are structurally weaker at design than at code because design quality requires human taste as part of the grading loop, and novelty is more important in design than in engineering.
    • The same feature released at different model capability levels can have completely opposite market outcomes — feature timing relative to model maturity is a first-class product variable.
    • Zone defense PM — spreading product people to cover gaps in a high-chaos, everyone-builds environment — is replacing roadmap ownership as the primary PM operating model at AI-native companies.
    • Eliminating the PM role in favor of 'everyone is a builder' discards the institutional knowledge encoded in the discipline — knowledge about what has been tried, failed, and why.
    • Dogfooding your own product for real work, even when it is suboptimal, compresses the discovery cycle to near-zero and generates the most honest product feedback available.
    • Long-horizon plans in AI product contexts should be intentionally hazy — any added precision beyond a short horizon is false precision that wastes time and creates misleading anchors.
    • Professional identity tied to specific tools or processes is fragile; identity tied to outcomes is durable. The gatekeeping layer of most roles — deep tool expertise — is eroding fastest.
    • The 'baby product' — a simplified codebase replica — is emerging as a concrete process primitive that replaces Figma interactive prototypes for interaction exploration in AI-capable teams.
    Action Items (8)
    • Audit your current product process against the new cost structure: identify which steps exist primarily to de-risk expensive implementation, and ask whether those steps still justify their overhead when implementation is cheap.
    • Establish explicit stage labeling for all artifacts (docs, prototypes, working code) — because medium no longer signals stage, your team needs a shared language for 'this is exploration' vs. 'this is a shipping candidate.'
    • Review your AI feature graveyard: list features that failed and assess whether the failure was a product-market fit failure or a model-readiness failure. Schedule a re-evaluation cadence tied to major model releases.
    • If your team does not have a 'baby product' equivalent, explore building one — a simplified codebase replica that allows rapid interaction exploration without touching production. Treat it as a replacement for Figma interactive prototypes.
    • Map where your product people are spending their time and identify coverage gaps — apply the zone defense mental model to find areas of the product where no one is providing steering or curation.
    • For any AI feature currently in development, explicitly document the model capability assumptions it depends on. This makes it easier to know when to re-test a shelved feature as models improve.
    • Separate your long-horizon intent (themes, bets, areas of interest) from your short-horizon plan (detailed, actionable). Stop adding false precision to anything beyond 6-8 weeks in an AI-driven product context.
    • Identify the parts of your current role that are tied to specific tools or process steps vs. the parts tied to outcomes and judgment. Invest deliberately in the latter — these are the parts that compound over model generations.
    Skip if: Skip if you are looking for tactical Codex usage tutorials or benchmarks — the episode is almost entirely conceptual and process-oriented, with only light coverage of specific Codex features.