.png)
A solo founder recently launched a scheduling app built entirely by describing features to an AI tool and accepting whatever code came back, never opening a single file to check it. The app worked, right up until a user found they could view another user's private calendar just by changing a number in the URL. Nobody had reviewed the authorization logic, because nobody involved in building the app had actually read any of the code.
That's the sharper edge of vibe coding, a term that went from a casual tweet to Collins Dictionary's Word of the Year in less than a year. Andrej Karpathy coined it in February 2025, describing a way of coding where you fully give in to the process and stop reading the diffs the AI produces. It's a genuinely useful way to move fast on a prototype. It's also, without a deliberate verification step added back in, a fast way to ship something you don't actually understand.
This piece explains what vibe coding is, the tools people use for it, the specific risks it introduces, and the concrete steps that turn AI-generated code from a gamble into something you can trust.
At its core, vibe coding means describing what you want in plain language, letting an AI tool generate the implementation, running it, and iterating through more prompts rather than by editing code directly. The person building the software treats the code itself as an implementation detail the AI owns, moving through a loop of describing a feature, running what comes back, and describing what's wrong.
That's meaningfully different from using an AI coding assistant as a fast typist while you still read and understand every change. Simon Willison, a widely cited voice on the topic, draws a sharp line here: someone who reads, tests, and understands everything the model writes is doing AI-assisted programming, not vibe coding. The distinction matters because it determines who's actually responsible for catching a mistake before it ships, and in true vibe coding, the answer is often nobody.
The origin matters because it explains the scope the term was meant to have. Karpathy described the workflow with disarming honesty, saying he accepts everything the AI suggests and no longer reads the diffs it produces. Crucially, he was describing weekend projects and throwaway experiments, not a production methodology.
The term escaped that original scope quickly. Within a year it had become a mainstream way to describe building real, shipped software, not just weekend experiments, which is exactly where the risks below start to matter far more than they did for Karpathy's original disposable side projects.
A range of vibe coding tools now sit at different points on the spectrum from "assistant" to "autopilot." Some are conversational coding agents built into an editor, where you describe a change and the tool edits files directly. Others are full application builders, browser-based tools where you describe an entire product and get a deployed, working app with minimal setup.
Agentic coding tools that can plan, write, run, and iterate on code across an entire project without a human editing files directly sit at the far end of that spectrum.
What most of these tools have in common is that they're capable of producing far more code, far faster, than any human reviewer can casually verify by reading it once. That capability is the whole appeal, and it's also exactly what makes verification necessary.
The risks aren't hypothetical, and they aren't really about the AI being "bad" at coding. Models can write functional-looking code that fails in specific, easy-to-miss ways:
Security vulnerabilities that pass a quick test. Code that works for a normal user can still leave an authorization check out entirely, the exact scenario in the introduction, where changing an ID in a URL exposed someone else's data.
Logic that works for the happy path only. A feature can look complete in a five-minute demo and still mishandle empty input, concurrent requests, or an edge case nobody thought to try.
Silent architectural drift. An AI tool asked for "one more feature" repeatedly can gradually restructure a codebase in ways that make sense locally but create inconsistency nobody notices until something breaks somewhere unrelated.
A false sense of understanding. Because the code runs and looks reasonable, it's easy to believe you understand a system you've never actually read, right up until you need to debug or extend it.
None of these are unique to AI-written code, human developers make the same mistakes. What's different is the volume and speed at which they can accumulate when nobody is reading the output as it's produced.
Verification doesn't mean abandoning the speed vibe coding offers, it means adding a deliberate checkpoint before code goes anywhere that matters.
Read the diff before accepting it, at minimum for anything touching authentication, payments, or user data. This is the single habit that separates vibe coding from the disciplined version of AI-assisted development.
Write or generate tests for the specific behavior you asked for, and check that at least one test would actually fail if the feature were broken. A test that just confirms whatever the code currently does proves nothing.
python
def test_user_cannot_view_others_calendar(client, other_user_event_id):
response = client.get(f"/events/{other_user_event_id}")
assert response.status_code == 403 # not 200
A test like this, written before or immediately after the feature, would have caught the scheduling app's authorization gap before a real user found it.
Run a dependency and security scan on anything before it reaches production, since AI tools can pull in outdated or vulnerable packages as readily as a human developer under deadline pressure can.
Ask the AI to explain its own code back to you, then check that explanation against what you actually see. If the explanation and the code don't match, that's a signal to look closer, not a reason to move on.
Picture reviewing AI-generated code for an endpoint that returns a user's calendar event:
python
@app.route("/events/<event_id>")
def get_event(event_id):
event = db.query(Event).filter_by(id=event_id).first()
return jsonify(event.to_dict())
This runs without errors and returns correct data for the user who created the event. It also returns correct data for anyone who guesses or increments an event ID, because nothing checks who's asking. A quick read catches this in seconds; a demo that only ever tests with the "right" user never surfaces it at all.
python
@app.route("/events/<event_id>")
def get_event(event_id):
event = db.query(Event).filter_by(id=event_id).first()
if event is None or event.user_id != current_user.id:
return jsonify({"error": "Not found"}), 404
return jsonify(event.to_dict())
The fix is small. Finding it required someone to actually look, which is the entire argument for verification as a non-negotiable step rather than an optional nicety.
|
S.No |
Aspect |
Vibe Coding |
Verified AI-Assisted Development |
|
1 |
Code review |
Rarely or never read |
Read and understood before merging |
|
2 |
Best suited for |
Prototypes, throwaway experiments |
Production features and anything user-facing |
|
3 |
Testing |
Often skipped or minimal |
Tests written against actual requirements |
|
4 |
Risk profile |
Higher, especially for security and edge cases |
Lower, catches issues before deployment |
|
5 |
Speed |
Fastest |
Slightly slower, but sustainable |
Treat vibe coding as a prototyping mode, not a production default, reserving the fully hands-off loop for throwaway experiments the way it was originally used. Add a review gate before anything reaches real users, even a fast one. Keep tests around authentication, payments, and data access non-negotiable regardless of how the code was written. And use AI Powered Playwright Automation to generate and maintain end-to-end tests alongside the feature itself, so verification scales at something closer to the same speed the code was generated.
Reading code critically, even code you didn't write, matters more now than it did before AI could generate functional-looking output instantly. Writing tests that actually exercise a requirement, not just confirm current behavior, catches far more than a visual read alone. Basic security instincts, knowing to check authorization on every endpoint that touches user data, close the exact gap in the scheduling app example. And comfort with automated testing tools, including AI-assisted ones like Playwright's newer agents, lets verification scale closer to the speed at which AI tools can generate code.
It's workable for a real product only with a verification layer added back in, code review, tests, and security checks before anything reaches users. Used exactly as originally described, no review at all, it's better suited to disposable prototypes than anything with real users or data.
The difference is whether you read and understand what the AI produces. Using an assistant while reviewing its output is a faster way to write code you still understand. Vibe coding, in its original sense, means not reading the code at all.
Security gaps that don't show up in casual testing, like missing authorization checks, are the most damaging because they're invisible until someone deliberately or accidentally exploits them. A non-technical founder without a technical reviewer is especially exposed to exactly this kind of risk.
Focus verification effort where it matters most, authentication, payments, and data access, rather than treating every line equally. A short, targeted test suite and a five-minute diff read catch most serious issues without eliminating the speed advantage.
Yes. It can generate and maintain end-to-end tests that exercise real user flows, including the unhappy paths a demo-focused build often skips, which is exactly where vibe-coded applications tend to have the most undiscovered problems.
Vibe coding isn't inherently reckless, it's a mode built for speed on disposable projects, and it does that job well. The risk shows up specifically when that mode gets applied to something people actually depend on, without adding back the review, testing, and security checks that Karpathy's original weekend-project framing never needed in the first place.
If you're building something real with heavy AI assistance, the practical takeaway isn't to slow down everywhere, it's to know exactly where to add a deliberate pause: authentication, payments, and anything touching another user's data.
Looking at your own most recent AI-generated feature honestly, did anyone actually read the diff before it shipped?
Follow NareshIT for more practical insights on technology, skills, and career development.