Category: Field Notes

  • Evaluation Is the New Taste

    Evaluation Is the New Taste

    A lot of people are still talking about AI like the main advantage is output.

    Write more posts. Generate more images. Summarize more research. Ship more code. Create more options.

    That is all useful. I use AI every day for exactly those kinds of things.

    But the more I build with it, the more convinced I am that output is not the real advantage anymore.

    Evaluation is.

    AI has made generation cheap. Not free, but cheap. You can get ten drafts, twenty logo concepts, a working prototype, a first pass at positioning, a test plan, a content calendar, a customer research summary, or a calorie estimate from a food photo in seconds.

    The bottleneck is no longer, “Can I make something?”

    The bottleneck is, “Is this any good?”

    That sounds simple, but I think it is one of the most important shifts in how modern builders need to operate.

    For a long time, taste mostly lived in someone’s head. A good founder, designer, product leader, engineer, or marketer developed intuition over time. They knew what felt right. They knew what was sharp, what was sloppy, what was overbuilt, what was confusing, what was close enough to ship, and what needed another pass.

    That kind of taste still matters. Maybe more than ever.

    But in an AI-native workflow, taste cannot only live in your head. You have to turn it into a system.

    That means examples.
    Rubrics.
    Golden sets.
    Review loops.
    Regression tests.
    Feedback cycles.
    Clear definitions of what “better” actually means.

    This is one of the lessons I keep running into across both my day job and what I am building with Mycro.

    At Psych Hub, we have been experimenting with AI workflows inside product and content operations. One example that stuck with me was an internal self-improvement pipeline. The basic idea was good. We had a way to create a heartbeat around improvement, surface opportunities, and keep the system moving.

    But the failure mode was obvious too.

    It was too eager to improve.

    That sounds like a good problem to have, but it is not. When a system is constantly looking for something to change, it can start making vibe improvements for the sake of making changes.

    A little clearer.
    A little more polished.
    A little more detailed.
    A little more “helpful.”

    But was it actually better?

    That is the question that matters.

    Without a scorecard, “improvement” becomes subjective in the worst way. The system starts rewarding movement instead of progress. It starts treating change as evidence of value.

    But sometimes the right answer is: nothing needs to change.

    That is a hard thing for AI systems, and honestly for product people, to accept. We like iteration. We like momentum. We like the feeling that we found something and made it better.

    But real evaluation has to create permission to leave good work alone.

    That was the missing piece. The self-improvement loop needed to be grounded in a scorecard. It needed a way to ask:

    What are we actually optimizing for?
    What would make this materially better?
    What would make it worse?
    What does “good enough” mean?
    When should the system recommend no change?

    That last question is underrated.

    A good evaluation system should not only identify what failed. It should protect what is already working.

    I have been seeing the same lesson while building Mycro.

    Mycro is a weight-loss app with AI woven through the product. One obvious place AI can help is calorie estimation. A user can describe a meal or take a photo, and the app can help estimate what they ate.

    That sounds simple until you actually care whether it works.

    Because “pretty good” is not good enough.

    If someone logs a grilled chicken sandwich from a major chain, the estimate should probably be close. If they log a homemade bowl with rice, ground beef, avocado, and sauce, the app needs to make a reasonable judgment call. If they only provide a photo, the system should know what it can and cannot infer. If they provide a photo and a description, it should use both. If the portion size is unclear, it should say that instead of pretending to know.

    The goal is not perfection. Humans are not perfect at calorie estimation either.

    The goal is trust.

    And trust does not come from vibes. It comes from repeated exposure to a system that behaves reasonably, explains itself well, and improves over time.

    That is where evaluation becomes the work.

    For Mycro, this is not just me eyeballing a few meal logs and deciding whether the estimates feel decent. I have been building a more structured evaluation harness around the kinds of meals real users actually log.

    Photos only.
    Descriptions only.
    Photos plus descriptions.
    Chain restaurant meals with published nutrition.
    Homemade meals weighed on a food scale.
    Easy cases.
    Ambiguous cases.
    Messy real-life cases.

    The point is not to create a perfect scientific instrument. The point is to stop relying on “that seems decent” as the evaluation method.

    Once you have a real evaluation set, even a small one, the way you build changes.

    You can compare models.
    You can test prompt changes.
    You can catch regressions.
    You can see whether a new approach is actually better or just more confident.
    You can measure where the system is weak.
    You can separate “this feels smarter” from “this performs better on the cases we care about.”

    That last part matters.

    AI systems are very good at sounding better. They are not always better.

    A response can be more polished and less accurate. More detailed and less useful. More confident and less honest. More impressive in a demo and worse in production.

    This is why I think evaluation is becoming one of the highest-leverage builder skills.

    The best builders are not just going to be the people who know how to prompt well. Prompting matters, but it is only one layer.

    The real advantage is knowing how to build the loop:

    Define the job.
    Generate the output.
    Evaluate the output.
    Capture what failed.
    Improve the system.
    Run it again.

    That loop applies everywhere.

    If you are using AI to write, you need to know what makes the writing good. Not just “make this punchier,” but what voice, structure, clarity, restraint, and audience fit actually mean.

    If you are using AI to code, you need to know what done means. Not just whether it compiles, but whether it is maintainable, secure, testable, and aligned with the product.

    If you are using AI for research, you need to know what a good answer requires. Sources, tradeoffs, uncertainty, recency, and whether the synthesis actually helps you make a decision.

    If you are using AI inside a product, you need to know what the user is trusting the system to do, and what happens when it gets it wrong.

    That is the part I think gets lost in a lot of AI conversation.

    People talk about speed like speed is the whole game.

    Speed is useful. I love speed. Speed is what makes it possible for a solo builder to do things that used to require a much larger team.

    But speed without evaluation just creates more noise.

    More drafts you do not ship.
    More features you do not understand.
    More content that sounds like everyone else.
    More prototypes that look impressive but do not survive contact with users.
    More “improvements” that quietly make the system worse.

    The real power is not moving faster in every direction.

    It is learning faster in the right direction.

    That requires judgment. And increasingly, it requires turning judgment into infrastructure.

    I think this is one of the big differences between casual AI usage and serious AI-native building.

    Casual AI usage asks, “Can this help me do the task?”

    Serious AI-native building asks, “How do I design a system where the work gets better every cycle?”

    That is a very different question.

    It moves you from prompting to orchestration.
    From output to feedback.
    From demos to durability.
    From taste as a personal instinct to taste as an operating system.

    And to be clear, I do not think this removes the human from the process. I think it makes human judgment more important.

    AI can generate options. It can compress time. It can help you move across disciplines. It can give you a first draft, a second opinion, a simulation, a critique, or a working prototype.

    But you still have to know what good looks like.

    You still have to decide what matters.
    You still have to notice when something is subtly wrong.
    You still have to understand the user, the context, and the consequences.
    You still have to make the call.

    The difference is that now, the best builders will not just make those calls manually every time. They will encode their judgment into repeatable systems.

    That might be a checklist.
    A benchmark.
    A content review rubric.
    A QA script.
    A customer feedback loop.
    A golden set.
    An eval harness.
    A scorecard.
    A workflow that forces the right questions before something ships.

    None of that sounds as exciting as “agents.”

    But I think it is where a lot of the real leverage lives.

    The future belongs less to people who can produce the most with AI, and more to people who can tell what is worth producing.

    Evaluation is the new taste.

    And the builders who take that seriously are going to compound faster than everyone else.

  • Stubbornness Is an AI Skill

    Stubbornness Is an AI Skill

    One of the most underrated skills in working with AI is stubbornness.

    Not prompting.

    Not knowing the name of every new model.

    Not having the perfect stack of agents, plugins, CLIs, MCP servers, and automation frameworks.

    Stubbornness.

    The willingness to keep going after the first output is bad.

    Because the first output is often bad.

    This is where I think a lot of people bounce off AI. They try it for something real, get a mediocre result, and immediately go back to the way they were doing it before.

    “AI can’t write good enough articles.”

    “It takes more time to review than to write it myself.”

    “The code is too buggy.”

    “It overcomplicates simple things.”

    “It gets too much wrong.”

    “It’s too generic.”

    “I tried it. It wasn’t useful.”

    And honestly, I get it.

    A lot of those complaints are true.

    AI does write generic first drafts. It does produce buggy code. It does miss edge cases. It does over-engineer simple features. It does confidently produce work that looks complete but falls apart under review.

    That is the part AI boosters often skip.

    Most of my AI journey has not felt like some clean, futuristic productivity montage. It has felt like fumbling through half-working tools, bad assumptions, broken dependencies, confusing docs, weird environment issues, and outputs that miss the mark in ways that are hard to explain.

    Even getting a basic environment running can be an ordeal.

    You are wiring up MCPs, skills, plugins, CLIs, local models, API keys, dependencies, permissions, repo conventions, and whatever else the current toolchain requires. You spend hours trying to get to one working sample app, one decent workflow, one useful content asset, or one internal tool that actually does what you need.

    It is not always elegant.

    Sometimes it is just brute force.

    I experienced this again recently while working on content marketing for Mycro.

    I am building a separate content repo to help with the heavy lifting: social assets, carousel concepts, short-form scripts, reusable content systems, and the pieces around them. The goal is not to have AI magically “do marketing.” The goal is to build a system that helps me produce better work more consistently without starting from zero every time.

    At first, the output was absolute shit.

    Generic. Overwritten. Too polished in the wrong ways. Missing the point. Saying things Mycro would never say. Creating assets that technically followed the prompt but clearly did not understand the taste, restraint, or positioning I wanted.

    It had the shape of content.

    It did not have the judgment.

    This is the moment where most people quit.

    And to be clear, quitting feels rational.

    If the output is bad, reviewing it takes a long time, and you still have to rewrite half of it yourself, it is easy to decide the whole thing is a waste of time.

    The same thing happens with writing.

    You ask AI for an article draft and it comes back sounding like every bland LinkedIn post you have ever scrolled past. The structure is predictable. The language is sanitized. The insight is obvious. It uses phrases you would never use.

    So the natural response is: I could have written this faster myself.

    And maybe you could have.

    The same thing happens with code.

    You ask AI to build a feature and it gives you something that technically kind of works, but the implementation is bloated. It creates abstractions you did not ask for. It ignores patterns already in the codebase. It misses obvious edge cases. It introduces bugs. It turns a simple fix into a small architecture project.

    So the natural response is: I could have done this faster myself.

    And again, maybe you could have.

    But that is also the trap.

    If you are measuring only the first output, AI often looks worse than doing the task manually.

    The article is not good enough. The code is not clean enough. The carousel is not sharp enough. The workflow is not stable enough. The automation breaks too easily. The agent needs too much babysitting.

    But the first version is not the verdict.

    The first version is diagnostic information.

    Why was the writing generic?

    What examples would have anchored the voice?

    What words would I never use?

    What point of view was missing?

    Why was the code bloated?

    What constraint would have forced a simpler implementation?

    What test would have caught the bug?

    What existing pattern should it have followed?

    Why did the content miss?

    Was the strategy unclear?

    Was the hook lazy?

    Was the CTA too thirsty?

    Was it trying to sound like a brand instead of a person?

    That is the loop.

    You are not just asking AI to make a thing.

    You are using the bad version to discover the missing instructions, missing examples, missing tests, missing constraints, and missing taste.

    Then you make those findings durable.

    That last part matters.

    The goal is not to win one prompt.

    The goal is to improve the system.

    Instead of accepting the bad output, you keep pressing.

    Why is this bad?

    What is missing?

    What would make this feel more like us?

    What would make the code simpler?

    What quality bar should this have hit?

    You take the bad output to another LLM entirely and ask for critique. You ask it to identify where the strategy is weak, where the voice is off, where the code is fragile, where the implementation is overbuilt, where the claims feel generic, where the asset loses the thread.

    Then you bring that feedback back into the builder.

    You do not just say, “Make it better.”

    You say, “Here is what failed. Here is why it failed. Here is what good looks like. Now update the skill so this does not happen again.”

    You turn taste into instructions.

    You turn recurring misses into QA checks.

    You turn vague preferences into examples and anti-examples.

    You turn coding mistakes into tests, lint rules, conventions, and review steps.

    You turn “this does not sound like me” into voice rules.

    You turn “this code is too much” into implementation constraints.

    You turn “this carousel is boring” into a sharper content rubric.

    You make Claude and Codex work in the same terminal session so they can check each other’s work. You let one model create, another critique, and another pressure-test the result against the actual goal.

    None of that happens from one prompt.

    It happens because you were annoyed enough to keep going.

    The leverage does not usually show up at the beginning. At the beginning, AI often sucks your time. You can spend a whole day trying to generate one reel, one carousel, one landing page section, one feature, one internal tool, or one workflow that would have been faster to do manually.

    That feels like failure if you are measuring only that one artifact.

    But sometimes the artifact is not the point.

    The point is the system you are building behind it.

    The first article takes too long because you are not just writing an article. You are finding your voice rules. You are building your editorial bar. You are teaching the system what claims are too generic, what examples matter, and what kind of language feels false.

    The first feature takes too long because you are not just shipping a feature. You are teaching the system your repo structure, your preferred patterns, your tolerance for abstraction, your testing expectations, and the kind of code you do not want to maintain later.

    The first carousel takes all day because you are not just making a carousel. You are defining the content strategy. You are shaping the voice. You are building the review loop. You are discovering what the AI gets wrong. You are creating the skill that will make the tenth carousel easier and the fiftieth carousel much better.

    That is the trade.

    AI work often feels inefficient right before it becomes leverage.

    It is slow, clumsy, and frustrating until the moment the system starts to hold. Then suddenly something changes. You are not starting from scratch anymore. You have a workflow. You have reusable context. You have QA. You have examples. You have tests. You have a machine that can help carry more of the load.

    That is when it starts to feel a little magical.

    Not because the model became perfect.

    Because you stayed with the problem long enough to teach the system what good looks like.

    I think this is the real dividing line for AI adoption.

    It is not between people who use AI and people who do not.

    It is between people who treat bad output as proof that AI does not work and people who treat bad output as diagnostic information.

    The first group tries it, gets frustrated, and goes back.

    The second group keeps asking:

    Why did this fail?

    What context was missing?

    What rule would have prevented it?

    What example would have helped?

    What test should have caught this?

    What should the next agent check before calling this done?

    That mindset compounds.

    Because once you start thinking this way, you stop seeing AI as a vending machine where prompts go in and finished work comes out.

    You start seeing it as rough but powerful system-building material.

    You still need judgment.

    You still need taste.

    You still need domain expertise.

    You still need to know when something is wrong.

    Maybe more than ever.

    But if you have those things and you are willing to be stubborn, the ceiling gets much higher.

    You can build workflows that would have required a team.

    You can create internal tools that would have sat in a backlog forever.

    You can turn repeated manual work into reusable systems.

    You can make multiple models critique, improve, and pressure-test each other.

    You can move from “AI helped me make this one thing” to “AI helped me build the machine that makes this kind of thing.”

    That is the shift.

    And it is uncomfortable because the early stages do not feel like leverage. They feel like wasted time.

    But that is true of a lot of meaningful skill acquisition.

    The first time you build the system, it is slower than doing the task yourself.

    The second time, it is still probably slower.

    Then one day, it is not.

    That is why stubbornness matters.

    You have to be willing to push through the awkward stage where the writing is not good enough yet, the code is not clean enough yet, the workflow is not stable yet, the output is not on brand yet, and the whole thing feels like it might be more trouble than it is worth.

    Sometimes it will be.

    Not every task deserves a system. Not every workflow needs AI. Not every bad output is worth saving.

    But the people who never push past the first bad version will never find the parts that do scale.

    They will keep saying AI is not useful for serious work.

    Meanwhile, the stubborn people will be quietly turning their frustration into infrastructure.

  • Don’t Sleep on Local and Specialized Models

    Don’t Sleep on Local and Specialized Models

    Don’t Sleep on Local and Specialized Models

    A lot of the AI conversation still revolves around the biggest cloud models.

    Claude.
    ChatGPT.
    Gemini.
    Codex.

    And to be clear, I get it.

    I still use Claude and Codex for most of my serious coding work. When I need deep reasoning, architecture help, debugging, planning, or a strong general-purpose collaborator, the frontier models are still where I spend most of my time.

    But I think there’s another part of the AI stack that people are underestimating.

    Local models.
    Specialized models.
    Narrow tools that are really good at one specific part of the workflow.

    Not because they are going to replace the frontier labs.

    They probably won’t.

    At least not broadly.

    OpenAI, Anthropic, Google, and others have enormous advantages. They have more compute, more research talent, more infrastructure, and more ability to push the edge of general intelligence.

    But in actual product workflows, that may not matter as much as people think.

    Because most products don’t need one model to do everything.

    They need the right model for the right job.

    And increasingly, I think the future of AI product development looks less like this:

    Pick the smartest model and send everything to it.

    And more like this:

    Use expensive frontier models for the hard thinking, then use local and specialized models for execution.

    That split is already starting to show up in my own workflows.

    The model is not the workflow

    When people talk about AI tools, they often talk about them like you have to pick one.

    Which coding tool do you use?

    Which image model do you use?

    Which voice model do you use?

    Which chatbot do you use?

    I think that’s the wrong framing.

    The model is not the workflow.

    The workflow is the system you build around the models.

    I still use Claude and Codex heavily for coding. They are great at reasoning through messy problems, shaping architecture, writing code, reviewing edge cases, and pushing through ambiguity.

    But once I get into content production, media generation, and repeatable execution work, the stack starts to look very different.

    Some models are better for transcription and timing.

    Some are better for voice.

    Some are better for image generation.

    Some are better for character consistency.

    Some are better for cleanup, background removal, or formatting.

    Some are better for animation or short-form video.

    And some are only worth using when the quality premium justifies the cost.

    That’s the important part.

    The goal is not to find one perfect AI tool.

    The goal is to build a workflow where each model does the thing it is best at.

    The Exact Stack Matters Less Than the System

    The specific tools are going to keep changing.

    That is part of the point.

    Six months from now, some of the models I use today will be better. Some will be cheaper. Some will be replaced entirely. Some new model nobody is paying attention to right now may become the obvious choice for one step of the process.

    So I don’t want to build a workflow that depends too heavily on one vendor, one model, or one assumption about what will be best.

    The architecture is the durable part:

    • Use frontier models where judgment matters.
    • Use local models where iteration speed matters.
    • Use specialized models where a narrow task has a better tool.
    • Use hosted APIs where the quality premium is worth the cost.
    • Build the workflow so pieces can be swapped out as better options emerge.

    That is the part I think more teams should be thinking about.

    The exact stack matters less than the philosophy behind it.

    Because six months from now, some of the specific tools will be better. Some will be cheaper. Some will be replaced entirely. Some new model nobody is paying attention to right now may become the obvious choice for one step of the process.

    That’s why I don’t want to build a workflow that depends too heavily on one vendor, one model, or one assumption about what will be best.

    I want a system that can evolve.

    Cost changes behavior

    For us, the target is repeatable but cost-efficient production.

    That matters a lot.

    If we are still iterating on scripts, I don’t want every draft, test, and small change to create a new premium rendering cost.

    Voice is a good example.

    If every voice generation runs through a high-end hosted API, it changes the creative process. You start locking scripts earlier. You become more careful before rendering. Every iteration has a cost attached to it.

    That might be fine at the final production stage.

    But it is not ideal during exploration.

    When local TTS is good enough for a big chunk of the workflow, the economics change.

    Now we can experiment more freely.

    We can tweak scripts.
    Regenerate lines.
    Try different pacing.
    Adjust tone.
    Test narration.
    Throw away bad versions.

    Without worrying that every small change creates another bill.

    That is a big deal.

    Not because local is always better.

    It isn’t.

    Sometimes the best model is still the expensive cloud model. Sometimes the quality gap matters. Sometimes the hosted API is more reliable, more scalable, or easier to operationalize.

    But sometimes “best” means something else.

    Sometimes best means cheap enough to use constantly.

    Sometimes it means fast enough to stay in the creative flow.

    Sometimes it means private enough to run close to your own data.

    Sometimes it means customizable enough for your specific use case.

    Sometimes it means good enough that the premium option stops being worth it for that step of the workflow.

    That’s where local and specialized models get interesting.

    Frontier models for thinking, specialized models for doing

    The pattern I keep coming back to is pretty simple:

    Use frontier models for thinking.

    Use specialized models for doing.

    That doesn’t mean the smaller models are dumb. It just means they don’t have to be world-class at everything.

    They need to be good at their part of the system.

    One model plans.

    Another generates.

    Another transcribes.

    Another cleans up.

    Another creates.

    Another removes.

    Another varies.

    Another reviews.

    Another packages.

    The frontier model may still be the brain.

    But the rest of the system may be made up of smaller, cheaper, narrower tools that are better suited for execution.

    That architecture also makes the whole system more flexible.

    As new tools emerge, you can swap them in.

    If a better TTS model shows up, replace that part.

    If a better image variation model appears, plug it in.

    If local video generation gets good enough, move that piece closer to your own infrastructure.

    If a hosted model becomes too expensive, route around it.

    If a niche model gets better at one task than the general-purpose model, use it.

    That is the real advantage.

    Not just using AI.

    Not just picking the “best” AI tool.

    But building a workflow that can evolve as the model landscape changes.

    Because this space is moving too fast to hard-code your entire process around one vendor, one model, or one assumption about what will be best six months from now.

    The coding version of this

    The next place I want to test this more seriously is coding.

    Right now, I still rely heavily on Claude and Codex for software work. But even there, I’m starting to wonder if the same pattern applies.

    Maybe the most expensive reasoning model doesn’t need to do every step.

    I already use heavier models for planning and architecture, then cheaper or faster models for more of the implementation.

    In Claude terms, Opus may do the thinking while Sonnet does more of the build work.

    The question is whether a local coding model can replace some of that implementation layer.

    Can Claude or Codex create the plan, then hand the scoped execution to a local model?

    Can the local model make the code changes, run the tests, fix the obvious issues, and get close enough to what I’d expect from Sonnet?

    I don’t know yet.

    But that’s the experiment.

    And if it works, the implications are meaningful.

    Because agentic development is not one prompt and one answer.

    It’s loops.

    Plan.
    Implement.
    Review.
    Fix.
    Test.
    Restart.
    Try again.

    If every loop runs through the most expensive model, the economics start to matter.

    But if the highest-cost model only handles the highest-judgment parts of the process, and local models handle more of the execution, the operating model changes.

    That could make agentic development cheaper, faster, and easier to scale.

    Not because local coding models are suddenly better than Claude or Codex at reasoning.

    But because they may not need to be.

    They may just need to be good enough at following a well-scoped plan.

    That is a very different bar.

    The best model depends on the job

    The teams that win probably won’t be the ones who use the most expensive model for everything.

    They’ll be the ones who understand the job to be done at each step, then choose the right model for that job.

    Sometimes that will be a frontier reasoning model.

    Sometimes it will be a coding model.

    Sometimes it will be a local model.

    Sometimes it will be a hosted API.

    Sometimes it will be an open-source tool.

    Sometimes it will be some weird little model that most people haven’t heard of yet, but that happens to be perfect for one part of the pipeline.

    That is where this gets exciting.

    The future of AI products probably isn’t one model doing everything.

    It’s a stack.

    Frontier models for thinking.

    Specialized models for doing.

    Local models where cost, speed, privacy, iteration, and customization matter.

    And a workflow flexible enough to keep swapping in better pieces as the space matures.

    That’s why I wouldn’t sleep on local models.

    They may never beat Claude or ChatGPT at general reasoning.

    They don’t have to.

    They just have to be good enough at the execution layer.

    And in a lot of real product workflows, the execution layer is where the volume is.

    That’s where the cost is.

    That’s where the iteration happens.

    That’s where the bottlenecks show up.

    That’s where customization matters.

    And that’s where a stitched-together system of specialized models may end up being far more powerful than one expensive model trying to do everything.

  • We Rebuilt the Platform. More Importantly, We Rebuilt How We Build.

    We Rebuilt the Platform. More Importantly, We Rebuilt How We Build.

    Four years ago, we built the platform Psych Hub needed.

    It worked. It grew. It supported real customers. It carried us a long way.

    But over time, something became obvious.

    The system we built was not designed for the future we were trying to create.

    The codebase was hard to reason about.
    The database needed to be rethought from the ground up.
    The user experience had years of friction layered into it.
    Our backlog was full of things we knew mattered, but could never quite justify prioritizing.

    And we are not a massive team.

    At some point, incremental improvement started to feel like the real risk.

    So we made a decision that felt, at the time, about an 8 out of 10 on the insanity scale.

    We stopped building on top of the old system.

    Not fully. We still supported the platform. We still handled major issues. We still had customers relying on us.

    But we stopped pretending the old foundation could stretch forever.

    We decided to rewrite the game mid-play.

    I told people: give us two weeks. If this is a bad idea, we will know quickly and go back.

    We never went back.

    Getting on the Tallest Lift

    When I was in middle school, I decided to learn to snowboard.

    Instead of starting on the bunny slope, my friends and I got on the tallest lift at the mountain.

    We all fell getting off.

    It took hours to get down that first run. Falling. Sliding. Getting up. Falling again.

    By day three, we were flying.

    That is what this rebuild felt like.

    We did not ease into it.
    We did not slowly refactor around the edges.
    We did not spend months creating a perfect plan before touching the product.

    We got on the biggest lift and committed to figuring it out on the way down.

    What We Actually Rebuilt

    In a matter of weeks, we migrated the platform we had spent years building into a completely new system.

    Not a cosmetic update.
    Not a fresh coat of paint.
    Not a rewrite for the sake of rewriting.

    A dramatically better user experience.
    A normalized database from the ground up.
    A new backend.
    A new frontend.
    A new back office.
    Multiple products and proof-of-concepts consolidated under one house.
    A year’s worth of customer feedback and backlog items addressed in one motion.

    The new platform brings together training, content, administration, enterprise workflows, and new AI-enabled capabilities in a way the old system never could.

    Things we thought we would never get to are now live.

    That is still a little surreal to say.

    But the bigger story is not only what we shipped.

    It is how we shipped it.

    The New Development Model

    This rebuild was made possible by an agent-native development pipeline we designed from scratch.

    Memory files live inside the repo and evolve with the product.
    Skills generate tickets and research requirements.
    Orchestration logic decides whether work needs a single agent or a larger sub-agent flow.
    Multiple models review implementation plans and code before we do.
    Code review happens before humans ever see the pull request.
    Agents write tests, update tests, and revise their own work.
    Infrastructure is managed as code.

    Humans are still deeply involved.

    We supervise.
    We refine.
    We challenge.
    We make judgment calls.
    We decide what good looks like.

    But we are not operating the same way anymore.

    We are not typing every line of code by hand.
    We are not spending weeks waiting for perfectly groomed tickets.
    We are not treating prototypes as disposable theater.

    The prototype is the spec.
    The review is the process.
    The loop is the operating model.

    Our retros now are not conversations about story points.

    They are conversations about workflow pain, agent performance, context quality, review quality, and where human judgment needs to be inserted earlier.

    We did not just rebuild the platform.

    We rebuilt how we build.

    Why It Had to Happen

    The uncomfortable truth was simple:

    With a small team and big ambition, the status quo was not going to get us there.

    Waiting would have been slower than rebuilding.

    The old architecture could not support the speed we needed. It could not support the quality bar we wanted. And it definitely could not support agents operating effectively inside it.

    If we wanted to build something much bigger than our team size suggested was possible, the foundation had to change.

    So we changed it.

    The Part I’m Most Proud Of

    I have been a primary contributor on this rebuild.

    Not because I am the only one capable.

    Because this new model makes it possible.

    I can move like a team.

    Research competitors.
    Spin up prototypes.
    Port them into real repos.
    Migrate data.
    Refine UX interactions.
    Test flows.
    Review implementation.
    Push the system until the product feels like a leap forward instead of another marginal improvement.

    Features that used to take weeks now take hours.

    Not because we are cutting corners.

    Because the machine removes friction.

    That is the part that is hard to fully explain until you experience it.

    A lot of product leaders talk about using AI to prototype. That is useful.

    But this is different.

    We are not just making clickable demos faster.

    We are building and shipping a real enterprise platform with agents as first-class contributors.

    That is a different chapter.

    What This Proves

    The platform is live.

    The migration is complete.

    The user experience is dramatically better. Features that sat in the backlog for months or years are now part of one cohesive system.

    And the team now has a new operating model for what comes next.

    That matters.

    Because the lesson here is not “AI can help you code faster.”

    That is too small.

    The lesson is that small teams can be much more powerful than they used to be if they are willing to rethink the entire system around how work gets done.

    Not just the tools.
    Not just the prompts.
    Not just the prototypes.

    The operating model.

    Small teams can be mighty when they build the machine around them.

    Four years ago, we built what we could with the tools and model we had.

    Today, we are building differently.

    And now that we have seen what is possible, there is no going back.

    When we got off that tallest lift in middle school, we fell. Everyone does the first time.

    But you do not remember the falls.

    You remember realizing you could ride.

    This rebuild felt impossible not long ago.

    Now it feels like the only way forward.

  • I Got Tired of Clicking “Continue”

    I Got Tired of Clicking “Continue”

    I didn’t set out to build a system that builds products.

    I just got tired of clicking “continue.”

    The first real shift didn’t happen with some big architectural decision. It happened because I was behind on a feature and decided to just vibe in Cursor and brute force it with AI.

    No process. No system. Just prompting and shipping.

    And it worked.

    That was the first “oh shit” moment.

    Not in a hype way. In a very practical way. I realized I could take something from idea to working code in hours instead of days. Not perfectly, but real enough to ship and iterate.

    That’s when we started taking it seriously.

    We moved into CLI-based workflows with Claude and began thinking about what a real process could look like if AI wasn’t just a helper in the IDE, but the thing actually doing the work.


    The Old Model

    Before this, we were operating like most product teams.

    Idea → spec → roadmap → sprint → feature.

    AI was there, but lightly. Mostly inside the IDE. Helping write code faster, not changing how we worked.

    It still required:

    • PRDs
    • sprint planning
    • story point estimation
    • backlog grooming
    • manual releases
    • a bunch of SaaS tools stitched together

    It worked. But it was slow.


    Where It Actually Broke

    The first version of this “new” workflow wasn’t a system.

    It was just us using AI more aggressively.

    We built out skills in Claude Code. Stored them in a repo. Refined them over time. Got to a point where we could consistently ship real code through it.

    But everything was still manual.

    I was the orchestrator.

    I was:

    • telling it to continue
    • reviewing changes locally
    • committing code
    • waiting for CodeRabbit feedback
    • going back to Claude to fix issues
    • repeating that loop multiple times

    At some point I realized I wasn’t building product anymore.

    I was managing a workflow.

    And most of that workflow was just me clicking “continue” and trying to keep multiple terminal tabs straight.


    Enter Dumb Eric

    That’s when we built the first version of the orchestrator.

    A deterministic system that could:

    • pull in tasks from Linear
    • execute steps in order
    • hand off between stages
    • pause when human input was needed
    • move forward automatically when it wasn’t

    We called it Dumb Eric because it wasn’t trying to be smart.

    It just ran the flow.

    That alone changed everything.


    The Next Problem

    Once the orchestrator was in place, a new gap showed up.

    The system worked, but it lacked judgment.

    It could execute steps, but it couldn’t reason about them well.

    So we layered an LLM on top.

    Now we had:

    • deterministic orchestration for structure
    • LLM reasoning for decision-making

    That combination turned out to be the real unlock.

    Not autonomous agents running wild.

    Not rigid workflows.

    The blend of both.


    Then It Started Improving Itself

    As we used the system more, it started breaking in predictable ways.

    Edge cases. Bugs. inefficiencies.

    Instead of fixing those manually, we built a self-improvement loop.

    The system could:

    • identify issues
    • create tickets for itself
    • propose improvements
    • implement changes

    At that point, something interesting happened.

    It started getting better without us directly touching it.


    Then I Became the Bottleneck Again

    The more it improved, the more suggestions it generated.

    And suddenly I was back in the loop.

    Reviewing every change. Approving every improvement.

    Different work. Same bottleneck.

    So we added auto-merge with confidence thresholds.

    If the system was confident enough, it could ship its own improvements.

    Now Dumb Eric updates himself.


    Somewhere Along the Way, This Stopped Being a Product Workflow

    At this point, what we had wasn’t just a better dev process.

    It was a system that could build product end to end.

    Features. Documentation. Knowledge base updates.

    And then we extended it to content.

    Our course creation pipeline now looks exactly like our software pipeline:

    • Linear task with human inputs
    • AI writes brief, storyboard, research
    • human approves
    • AI writes scripts
    • human approves
    • AI produces content (voice, visuals, slides)
    • changes are made by prompting and re-rendering

    The stack is a mix of local models, animation tools like Rive, and code-based video editing with Remotion.

    But the important part isn’t the tools.

    It’s the workflow.


    What Disappeared

    We didn’t replace our old process with a better version.

    We made most of it unnecessary.

    No more:

    • PRDs
    • long fantasy roadmaps
    • sprint planning
    • story points
    • backlog grooming
    • manual releases
    • buying SaaS for everything
    • costly, manual course production workflows

    Course production used to be a real investment.

    Planning, scripting, recording, editing, revisions. It added up quickly in both time and cost.

    Now it runs through the same system as everything else.

    A course goes from idea → brief → script → production through a pipeline, with human approvals in the right places.

    The marginal cost of producing a course is basically zero.

    Not because the work disappeared.

    Because the system does it.


    What Surprised Me

    A few things I didn’t expect:

    Deterministic orchestration + LLM reasoning is far more effective than either alone.

    Agents don’t need to be that smart if the workflow is good.

    Self-improving systems actually work in practice.

    And the biggest one:

    The bottleneck isn’t building anymore.

    It’s human decision speed.


    What This Actually Is

    If I had to describe it in one sentence:

    It’s a system that builds product end to end.

    Right now, it’s a product factory.

    By the end of the year, it’ll be a company factory.


    How This Actually Happened

    This wasn’t designed upfront.

    It emerged.

    Every step came from removing pain:

    • too slow → use AI
    • too manual → build orchestrator
    • too rigid → add LLM reasoning
    • too fragile → add self-improvement
    • too dependent on me → add auto-merge

    Every time I became the bottleneck, I removed myself.


    Where to Start

    Don’t start by trying to build a system.

    Start by feeling the pain.

    Build something real first.

    For us, that meant:

    • using Claude Code skills
    • getting to a point where we could actually ship code through it
    • refining that process until it worked

    Only then did we start removing friction.

    If you don’t feel the pain, you’ll overbuild the system.


    The Shift

    The goal isn’t to use AI to build faster.
    It’s to build a system where product is the output, not the work.

    That’s the difference.

    Most teams are still trying to use AI inside their existing process.

    The real shift is building a system where the process disappears.

    Once you cross that line, everything changes.

  • The Blurred Lines

    The Blurred Lines

    Lately I’ve noticed something happening on our team that would have felt strange a few years ago.

    Me, our CTO, and our Director of Engineering are all doing essentially the same thing.

    Not identical work. Different strengths. Different lenses. But overlapping execution in a very real way.

    Steve, our CTO, is deep in the weeds testing tools, pushing on new skills, and experimenting with fully autonomous agents. I’m focused more on product feel, interactions, and how things actually come together in the experience. Bryan is dialed in on security, consistency, and quality.

    But we’re all pushing code.

    And the truth is, I don’t think we could be shipping nearly as fast right now if we tried to keep the lines clean.

    This would have been weird before

    Historically, there were clear boundaries for a reason.

    Product defined the problem.
    Engineering built the solution.
    CTO focused on architecture, systems, and long-term direction.

    And honestly, there was a time when engineering not wanting product anywhere near the code made total sense. It protected quality. It protected consistency. It kept ownership clear.

    But AI-assisted development is changing the cost of contribution.

    Now a product leader can meaningfully contribute to real implementation. Engineers are shaping product decisions in real time, not just reacting to specs. The CTO is in the tools every day, testing what’s possible and pushing the edge of how we build.

    The distance between idea and execution is collapsing.

    And with that, the old boundaries start to get in the way.

    We’re faster because the lines are blurred

    The biggest realization for me has been this:

    We could not be moving at this pace if we tried to preserve traditional role separation.

    Agentic development thrives on momentum. It rewards people who can see a problem, take a swing at it, and move it forward without waiting for a handoff.

    If every idea has to travel through a clean chain of:
    Product → Spec → Engineering → Review → Ship

    You lose time. You lose context. You lose energy.

    But when the same group of people can:

    • spot the problem
    • sketch the solution
    • try something
    • refine it
    • and ship it

    You compress weeks into days.

    That’s what we’re seeing right now.

    It takes trust. And it’s uncomfortable.

    This only works because there’s a lot of trust.

    We’ve worked together across multiple companies for 15+ years. We know how each other thinks. We know each other’s strengths. We know when to push and when to step back.

    And if I’m being honest, it takes some uncomfortable acceptance. Especially on the engineering side.

    Letting a product person contribute to the codebase isn’t a small shift. It challenges old instincts. It can feel risky. It can feel messy.

    But the reality is, the best ideas have always come from when the lines blur a bit.

    At Wondr, one of our engineers, Chase, had ideas that directly improved mobile app adoption. Not theoretical improvements. Measurable impact. Those ideas didn’t come from a spec. They came from being close to the problem and thinking like a product person while building like an engineer.

    That kind of cross-pollination isn’t new.

    What’s new is that now we’re not just sharing goals.

    We’re sharing execution.

    This isn’t for everyone

    I don’t think this model works everywhere. At least not yet.

    It probably breaks down in:

    • Large organizations
    • Low-trust environments
    • Teams that haven’t worked together long
    • Places without strong systems and safeguards

    If you don’t have maturity, clear workflows, and mutual respect, blurred lines can quickly turn into chaos.

    But for small, experienced, high-trust teams, the upside is massive.

    You get:

    • More velocity
    • Faster learning loops
    • Better ideas
    • More ownership
    • Less waiting

    And if you have the right guardrails in place around security, quality, and consistency, you can move fast without breaking everything.

    The future feels different

    This is one of the bigger shifts I’m seeing up close.

    It’s not just that AI helps engineers code faster.

    It’s that the definition of who builds is starting to change.

    Product leaders can execute.
    Engineers can shape product in real time.
    CTOs are experimenting directly in the tools.

    The roles don’t disappear. The strengths still matter. The lenses are still different.

    But the lines between them are getting harder to see.

    And in the right environment, that’s a good thing.

    The best small teams aren’t just aligned on goals anymore.

    They’re aligned in execution.

  • Small Teams Don’t Die From Bad Ideas. They Die From Slow Ones.

    Small Teams Don’t Die From Bad Ideas. They Die From Slow Ones.

    I’ve seen more teams struggle from moving too slowly than from making the wrong call.

    Not because they were lazy.
    Not because they didn’t care.
    Usually the opposite.

    They cared so much that everything had to be right before anything could ship.

    The problem is that early-stage companies don’t have the luxury of certainty. You have a runway. You have limited time. You need to learn fast enough to survive. If you take a year to carefully validate every decision, you may not have a company left by the time you’re confident.

    Bias toward action is not a personality trait. It’s a survival skill.


    The startup that never got the chance to learn

    I once started a company with a group of 6 cofounders. Half of us were product and tech. The other half were responsible for the program and content. The idea was a mobile app focused on stress and emotional health. A digital program with video content. Something that could genuinely help people.

    In the early days, it was great. We were forming the business, building mockups, getting the first version of the program off the ground. It felt like momentum.

    Then the pace started to shift.

    The content side became deeply focused on getting everything exactly right. Scripts were reviewed and rewritten over and over. Research was double checked. I remember one person who looked like they had been up for days, buried in papers, making sure every claim could be defended.

    It came from a good place. They cared about the quality. They wanted to stand behind it. They viewed this as releasing something deeply personal into the world.

    But it slowed us down in a way that started to matter.

    Getting the v1 content created took 4 to 5 months. When it was finally done, it felt like the finish line to them. I kept saying, this is the starting line. We’re about to learn all the things we need to change. We need to become a content factory while we’re a product factory.

    At the same time, we started debating smaller and smaller things. The shade of red in the design. The icon used on a screen. We would get aligned, then a week later the topic would come back up again. It was as if the product had to be perfect before it could exist.

    To be fair, the product side wasn’t perfect either. We took too long building the app. This was one of the first times I leaned heavily on AI to speed up development because we just needed to get something working. We could not spend months polishing something that hadn’t even met a real user yet.

    Eventually we got to the point where we shared the product with friends and family. And then we shut it down.

    Not because the idea was bad. Not because we ran out of money. We just couldn’t agree on how to operate.

    I wanted to move fast, learn from real users, take criticism, and iterate. The other side wanted to feel fully confident before releasing anything. They wanted it to be something they could stand behind completely.

    Both perspectives were understandable. But they were incompatible.

    It was either going to break before revenue or explode after it. We chose to stop early.

    Looking back, I still believe the idea had potential. And if we were building it today, the app could be delivered in record time. The program could be produced at a high quality in record time. The tools are better now. The speed is there.

    But speed only helps if the culture supports learning through motion.


    What bias toward action actually means

    Bias toward action is not recklessness.

    It’s not betting the farm.
    It’s not skipping thinking.
    It’s not ignoring risk.

    It’s making decisions without 100 percent information.
    It’s placing small bets.
    It’s learning through motion.
    It’s accepting that small mistakes are part of the process.

    You can move fast and still be responsible. You can create offramps. You can mitigate risk. You can test in controlled ways.

    The key is to stop waiting for perfect clarity before you start.

    Clarity comes from doing.


    Time is the real constraint

    Early stage teams don’t fail because every idea was wrong. They fail because they ran out of time before they learned enough.

    You have a runway.
    You have limited shots on goal.
    Your metrics are not where they need to be yet.
    Your next round is not guaranteed.

    If you take your time to feel confident in every decision, you may not have a company left to worry about.

    That sounds dramatic, but it’s the reality. Startups are a race against time, not a quest for perfection.

    And the pace of building and testing ideas is only accelerating. The cost to prototype, launch, and iterate has collapsed. The teams that learn the fastest will win. The teams that wait for certainty will fall behind.


    What slow cultures feel like

    You can feel a low-action culture almost immediately.

    Every decision needs approval.
    Small changes turn into meetings.
    Topics get revisited again and again.
    People start waiting instead of acting.

    I’ve seen environments where smart, capable people hesitate to change a single word on a web page because they’re worried about getting questioned later. Over time, that kills initiative. People stop thinking like owners. They start thinking like operators.

    Overthinking strips away autonomy. And without autonomy, you don’t get momentum.

    You get a lot of discussion.
    You get a lot of planning.
    You don’t get much movement.


    What high-action cultures look like

    I saw the opposite early in my career at Koddi. I was employee number one. We were bootstrapped. It was me, the president, and a small development team.

    We did everything. Product. QA. Sales. Customer support. Whatever needed to get done.

    We were relentlessly focused on moving fast with purpose. It wasn’t chaotic. It was disciplined. We removed distractions. We obsessed over making customers love us. We responded quickly. We solved problems quickly. We looked for the next opportunity before anyone asked.

    We grew fast. We stayed profitable. It felt like Seal Team training in execution.

    Later on, at Wondr, I saw how speed creates clarity.

    We saw data that showed mobile app users had significantly better weight loss success and adherence. That directly tied to revenue. Because we had autonomy, we shifted focus immediately. We pushed hard on mobile adoption and made a real impact in a short amount of time.

    At a lot of companies, that insight would have been recorded, discussed, prioritized, and maybe addressed months later. We just acted.

    Because we had autonomy, we shifted immediately. Speed turned insight into revenue.

    Plans matter. Roadmaps matter. But smart people need the space to break the plan when reality changes.


    Practical ways to build a bias toward action

    This isn’t about slogans. It’s about behaviors.

    Default to trying.
    Run small experiments.
    Shorten decision loops.
    Kill zombie discussions that keep resurfacing.

    You don’t need to solve everything at once. You just need to learn faster than the problems are growing.


    You will never feel fully ready

    There will always be one more thing to refine. One more thing to validate. One more expert opinion to get.

    If you wait until everyone feels completely confident, you will wait forever.

    Bias toward action doesn’t mean you don’t care about quality. It means you care about learning. It means you accept that the first version will be imperfect. It means you trust that you can improve once reality starts pushing back.

    Speed creates clarity.
    Action reveals truth.
    Momentum builds belief.

    Small teams rarely die from bad ideas.
    They die from not moving fast enough to find the right one.

  • Why I’m Less Afraid of AI Breaking Things Than Slow Teams

    Why I’m Less Afraid of AI Breaking Things Than Slow Teams

    Last week, an AI agent deleted our staging database.

    Not corrupted it. Not partially broke it. Deleted it.

    To its credit, it immediately owned the mistake. No hedging. No confusion. Just a clear apology and a summary of what happened.

    The good news? We had staging fully restored in about 20 minutes.

    That’s the part that stuck with me – not the failure, the recovery.

    The real lesson wasn’t that AI can make big mistakes. We already know that. The lesson was how little damage it actually caused in a team designed to move quickly.

    And it reinforced something I’ve been thinking about for a while:

    I’m less afraid of AI breaking things than I am of slow teams.

    Breakage is inevitable. Slowness is fatal.


    This isn’t an argument for recklessness

    Let’s be clear. Speed without guardrails is chaos.

    If you’re going to let AI operate in your repo, you need containment: proper environment separation, scoped permissions, protections against destructive commands, version control discipline, backups you’ve actually tested, and visibility into what’s being executed.

    That staging incident didn’t hit production – and that wasn’t accidental. That’s guardrail design.

    The goal isn’t “let it break things.” The goal is to reduce blast radius when it does. Because something eventually will.


    Software has always broken

    Humans push bad code. Servers go down. Migrations go sideways. Someone drops a table.

    None of that is new.

    What’s different now is speed. AI can investigate, modify, and execute faster than any junior engineer – sometimes faster than a senior one. Yes, that includes making mistakes faster. But it also includes fixing them faster.

    If your team can detect, diagnose, and recover quickly, most mistakes become small events. Annoying, but survivable. Sometimes even valuable.

    If your team moves slowly, even small problems turn into drawn-out, expensive disasters.

    The risk isn’t just that something breaks. The risk is that you can’t respond when it does.


    The real safety net is recovery speed

    We like to think safety comes purely from prevention: more approvals, more documentation, more review layers, more caution.

    Process matters. Especially in healthcare. Especially when people are involved.

    But process without recovery capability creates fragility, not safety.

    The older I get, the more I believe the real safety net is recovery speed.

    Do you have backups? Can you restore quickly? Can your team jump in and solve the problem without a week of meetings? Can you move forward again the same day?

    If the answer is yes, a lot of scary things become manageable.

    That staging incident could have been a nightmare if we didn’t have our fundamentals in place. Instead, it was a 20-minute disruption and a useful forcing function.

    That’s not luck. That’s posture.


    AI just exposes what was already true

    AI didn’t create this dynamic. It just made it more obvious.

    Fast teams have always had an advantage. Now the gap is widening.

    If you can prototype quickly, test quickly, recover quickly, and iterate quickly, you can afford to take more swings. You learn faster. You adapt faster. You improve faster.

    If you can’t, every mistake feels existential. So you slow down. You overanalyze. You try to control everything.

    And that’s where the real danger lives – not in breakage, but in paralysis.


    Slow teams don’t feel slow. They feel careful.

    This is the tricky part.

    No one thinks they’re moving slowly. It feels responsible. It feels thoughtful. It feels like protecting the company.

    I’ve been part of teams that spent weeks planning something that could have been tested in a day — long docs, endless alignment, multiple rounds of discussion, careful sequencing.

    By the time we shipped, the world had already moved on.

    In a startup, that kills you quietly.

    You don’t lose because of one catastrophic mistake. You lose because you learn too slowly.


    My fear has shifted

    A year ago, the idea of an AI autonomously making changes to a repo would have made me uneasy.

    Now? I’m still cautious. But I’m not nearly as afraid.

    Because I’ve seen what happens when you pair speed with guardrails: staging environments, backups, version control, clear ownership, tight feedback loops, and limited blast radius.

    When those things are in place, most mistakes are recoverable.

    What worries me more is the opposite environment – weeks to make a decision, months to ship something small, fear of touching anything, endless discussion with little movement.

    That’s the kind of system that slowly suffocates a company.


    The bar is changing

    We’re entering a phase where speed and quality aren’t the same tradeoff they used to be.

    AI can help investigate bugs in minutes. It can reason through codebases quickly. It can draft, test, and refine faster than we’re used to.

    But all of that only matters if the team is willing – and structurally able – to move.

    The teams that win won’t be the ones that avoid every mistake.

    They’ll be the ones that detect issues quickly, contain them, recover quickly, learn quickly, and keep going.


    I’d rather recover fast than move slow

    That staging incident didn’t make me want to lock everything down.

    It reinforced something I already believed.

    If you’re letting AI operate in your environment, you should absolutely be thoughtful. You should design guardrails. You should respect the risks.

    But you should also build a team and a system that can take a hit and keep moving.

    Breakage will happen – with humans, with AI, with both.

    The companies that survive won’t be the ones that prevent every failure.

    They’ll be the ones that design their systems so failures are contained – and recover faster than anyone else.

  • Prototyping > Vibe Coding

    Prototyping > Vibe Coding

    I’m not a fan of the term vibe coding. It sounds sloppy, unserious, and a little too close to “just prompt until something happens.” But the underlying idea, compressing the distance between an insight and something real you can interact with, is one of the most important shifts happening in product right now.

    For years, the path from idea to reality was slow and structured.

    Insight → wireframes → mockups → revisions → stakeholder input → engineering handoff → first working version → realization that we missed a lot.

    That cycle could take weeks. Sometimes months. And it often meant we were making major decisions based on static artifacts instead of something people could actually use.

    Now we can collapse most of that into hours.

    Tools like Lovable, Replit, and v0 let you spin up working prototypes in minutes. You can interact with them, tweak them, rethink flows, and explore directions before anyone commits to building production code. While these tools can generate full applications, I think their real sweet spot today is helping product teams and stakeholders spec projects faster and more clearly than ever before.

    Instead of describing the product, you can just show it.

    And more importantly, you can use it.

    From Spec to Something Real

    On our current platform project, we built a comprehensive prototype of the future product in about 20 hours and roughly $200. A few years ago that would have taken months and $30–40k in external design and engineering support just to reach the same level of clarity.

    But the part that really changed things wasn’t just the speed.

    We synced that prototype to GitHub. Then we had agents port the prototype pages into our actual codebase. At that point, the conversation shifted from:

    “Here’s a Figma file. Go build this.”

    to:

    “Here’s the UI. Your job now is to make this page work.”

    That’s a completely different starting point.

    The prototype becomes the spec.

    It already contains layout, interactions, component structure, and intent. Agents building against it already understand how things are supposed to behave. This is especially effective if your production stack shares the same front-end patterns. In our case, we chose building the with React, Tailwind, and shadcn because that is usually the default ecosystem most of these prototyping tools generate against. That alignment makes the handoff from prototype to production much smoother.

    The Tradeoff: Sameness

    Of course, there are tradeoffs.

    You can already see a certain sameness creeping into modern apps. Many AI-built interfaces use the same component libraries with slight visual variation. Dashboards start to look familiar. Patterns repeat.

    That’s the cost of speed.

    But it’s also the reason agents can move so fast. Standardized stacks mean less friction, less guesswork, and more momentum.

    And honestly, in early stages, speed matters more than visual originality. You’re trying to find something that works. Something useful, clear, and valuable. You can differentiate later.

    Why These Tools Matter Right Now

    Right now, tools like Lovable and Replit are incredible for getting an idea off the ground and iterating until something feels right. Not perfect. Not finished. Just right enough to build.

    One interesting thing I’ve noticed in my own workflow is how prompting style affects outcomes.

    Early on, I was extremely prescriptive. I’d use ChatGPT or Claude to generate long, detailed markdown instructions describing exactly what to build, then feed that into the prototyping tool. That works. Sometimes it’s necessary, especially when you have a clear vision.

    But I’ve also found value in being intentionally vague.

    Prompts like “make the dashboard useful” or “this page feels empty” can lead to unexpected ideas. Layouts, features, or small touches we hadn’t thought of. Occasionally, those end up being the best parts.

    There’s a balance there. Direction matters. But leaving room for interpretation can surface new thinking.

    The Real Shift

    To me, that’s what modern prototyping really is.

    It’s not about replacing design.
    It’s not about skipping engineering.
    And it’s definitely not about “vibes.”

    It’s about moving the moment of truth earlier.

    Instead of debating ideas, you interact with them.
    Instead of writing long specs, you explore working versions.
    Instead of waiting weeks to learn what you missed, you find out the same day.

    And once you have something real, agents can take it the rest of the way.

  • Why Agile Breaks in the Agent Era

    Why Agile Breaks in the Agent Era

    Agile was one of the most important shifts in how software teams work. It replaced rigid planning with iteration, feedback, and collaboration. It helped teams ship faster, learn sooner, and waste less effort.

    But Agile was designed around a core assumption:

    Humans do most of the execution.

    That assumption is now false.


    The assumption Agile is built on

    Agile practices – sprints, story points, backlog grooming, standups, all exist to solve a specific problem:

    How do we coordinate groups of humans doing complex, expensive work?

    When humans write code:

    • execution is slow
    • changes are costly
    • mistakes compound
    • coordination is necessary

    Agile optimizes around that reality.
    Planning reduces waste.
    Small increments reduce risk.
    Commitments help teams align.

    All of that makes sense—if humans are the bottleneck.


    What changed

    AI agents can now:

    • implement complex features
    • refactor codebases
    • write and fix tests
    • iterate rapidly with minimal cost

    Execution is no longer scarce.

    The constraint has moved.

    The hardest part of building software today is not typing code – it’s:

    • deciding what to build
    • deciding what quality looks like
    • deciding when something is “good enough”
    • deciding what to do next based on what you see

    In other words: judgment.

    Agile does not optimize for judgment. It optimizes for coordination.


    When execution becomes cheap, planning becomes noise

    In an agent-driven world:

    • estimating work is guesswork
    • sprint commitments become artificial constraints
    • backlogs grow faster than they’re resolved
    • teams plan more than they learn

    You can feel this tension already.

    Teams say they “do Agile,” but:

    • they skip estimates
    • they ship outside sprint boundaries
    • they prototype first and plan later
    • they use AI tools to bypass the process

    Agile hasn’t failed. Its assumptions have expired.


    The new bottleneck is review, not execution

    When agents can produce working software quickly, risk shifts downstream.

    The question is no longer:

    “Can we build this?”

    It’s:

    “Is this the right thing, at the right quality, at the right time?”

    That question can’t be answered in planning meetings.

    It can only be answered by:

    • seeing working software
    • interacting with it
    • judging it in context
    • learning from real signals

    This is why review beats planning in the agent era.


    Why sprints stop making sense

    Sprints exist to batch work into predictable intervals.

    But when:

    • execution is fast
    • changes are cheap
    • learning happens continuously

    Fixed cadences become friction.

    Work doesn’t finish because the sprint ends.
    It finishes when a decision is made.

    In practice, teams already know this—they ship when things are ready and pretend the sprint mattered.


    Agile optimizes throughput. The agent era demands judgment.

    Agile measures:

    • velocity
    • story completion
    • predictability

    But those metrics assume:

    • humans are the workers
    • output is the goal
    • efficiency is the constraint

    In agent-first teams:

    • output is abundant
    • efficiency is cheap
    • bad decisions are the real cost

    Optimizing throughput without improving judgment just means shipping the wrong thing faster.


    This isn’t about abandoning Agile values

    Agile’s values like collaboration, adaptability, and customer focus still matter.

    What breaks is the mechanism.

    The tools, rituals, and language of Agile reflect a world where:

    • execution is expensive
    • planning is protection
    • small increments are mandatory

    That world is disappearing – And it’s happening fast.


    What replaces it

    In the agent era, teams need a system that:

    • treats execution as cheap
    • treats judgment as scarce
    • prioritizes review over prediction
    • measures learning over commitment
    • adapts rigor based on risk

    This is why new models are emerging – whether teams name them or not.

    Some are already working this way.
    Most just don’t have language for it yet.


    The transition is already happening

    You can see it in:

    • prototype-first product work
    • AI-driven implementation
    • features built in single passes
    • planning rituals quietly abandoned
    • review and iteration happening in real time

    The gap is not practice.
    It’s the operating model.


    A final thought

    Agile replaced Waterfall because the world changed.

    The agent era is another such moment.

    When execution is cheap,
    planning loses power,
    and judgment becomes everything.

    The teams that recognize this early will build faster, better, and with fewer compromises.

    The rest will keep coordinating work that no longer needs coordination.