Home / Labs / Security

The security risks of vibe coding, and how to spot them before your customers do

An app that works is not the same as an app that is safe. Here are the nine security problems that show up again and again in AI-built software, in plain language, and what to do about each.

A brass padlock and a steel chain lying on a laptop keyboard under red and green light
Key takeaways
  • AI tools write code that runs, not necessarily code that is safe. A feature that works in a demo can still let one customer read another customer's data.
  • The same nine problems appear again and again: secrets left in the code, missing permission checks, unchecked input, invented software packages, security checks that do nothing, exposed admin pages, hidden instructions, unsafe file handling and careless use of personal data.
  • You do not need to be technical to ask the right questions. You do need someone independent to test the running application before real customers use it.
  • Often the problems can be fixed one by one. When permissions are scattered everywhere, rebuilding that layer is cheaper than patching it for ever.

The short answer

Is it safe to launch an app that an AI tool wrote from a few prompts? Not on the strength of the demo alone. "Vibe coding", the habit of describing what you want in everyday words and letting an AI assistant produce the software, is a wonderful way to get a first version quickly. It is also a way to collect security problems that nobody decided to create and nobody has looked for.

This article explains why those problems appear, describes the nine we meet most often, and ends with a practical way to decide whether to repair or rebuild. We build software with AI assistants ourselves, every week, so this is not an argument against using them. It is an argument for using them with your eyes open.

What "vibe coding" means, and where it is perfectly fine

The term describes a way of working in which the person asks an AI assistant for a feature in plain language, runs the result, and keeps asking for changes until it looks right. Nobody necessarily reads the code that comes out. The feedback comes from the screen: does the button work, does the page look as expected?

For a throwaway prototype, an internal experiment or a page that only you will ever use, that is a reasonable trade. The cost of a mistake is small, and speed is worth more than polish. The trouble starts when the prototype quietly becomes the product: someone adds a login, a few customers sign up, a payment form appears, and the original "quick experiment" is now holding real names, real email addresses and real money.

Nothing about that journey forces a security review. The app never stopped working, so no alarm ever went off. That is the central problem, and everything below follows from it.

Why AI-written code so often has security gaps

AI models are trained to produce code that does what you asked. "Allow users to download their invoices" is a clear, testable request. "Make sure no user can ever download somebody else's invoice" is a different kind of request: it describes something that must not happen, and a demo never shows what does not happen. Unless somebody states the rule, there is no reason for the code to enforce it.

Several habits of the way people use these tools make this worse:

  • Security decisions get lost along the way. In a long conversation with an assistant, an early instruction such as "only the owner may edit a record" can quietly disappear when a later change rewrites the same part of the code. Nobody notices because the feature still works.
  • Changes get too big to read. An assistant can produce hundreds of lines in seconds. A reviewer faced with that much new code tends to skim it, and a skimmed review is not a review.
  • Trust builds up. After ten changes that worked, the eleventh gets less attention. People relax precisely when they should be checking.
  • The assistant reviews its own work. Asking the tool "is this secure?" mostly returns reassurance, because it shares the blind spots of whoever wrote the code, which in this case is itself.
  • Apps appear outside the normal channels. A person in marketing or finance can now build a working tool in an afternoon and connect it to company data without telling anyone in IT.

One widely quoted industry study published in 2025 tested a large number of code samples written by AI models and found that a substantial share, close to half in that study, contained a known class of security weakness even though the code worked. Figures vary from study to study and move as the models improve, so treat any number as a rough indication and check the source before quoting it. The pattern, however, is consistent: working and safe are two separate properties.

The nine problems we see most often

Each problem below is described in everyday terms first, then with a short note on how it happens and how it is found. They are ordered roughly by how often they cause real damage.

1. Secrets left where anyone can read them

Every app uses secrets: the password to the database, the key that lets it send email, the token that connects it to a payment service. Assistants often put them directly in the code, because that is the shortest path to something that runs. Once the code is copied to a shared repository, a server or a browser, so are the secrets.

The dangerous version is a secret that ends up in the part of the app that is sent to every visitor. Anyone who opens the browser's developer tools can read it. Automated programs scan the internet around the clock looking for exactly these keys, and a leaked key can be used within minutes. Finding the problem is easy for a specialist and quick with free scanning tools. Fixing it takes more than deleting the line: any key that was ever exposed must be replaced, because it stays in the history of the code.

2. Missing permission checks

This is the most serious problem on the list, and the hardest to see. Logging in proves who you are. It says nothing about what you are allowed to do. An app can handle the first part perfectly and forget the second: if the address of your invoice is a number, and changing that number shows you somebody else's invoice, the app has an access-control failure.

Assistants drop these checks easily. Each screen and each action needs its own question, "does this person own this record?", and when code is rewritten many times, some of the questions disappear. A demo with one test user never reveals it. It shows up only when you test with two users, or with a user and an administrator, and try to cross the line on purpose.

3. Trusting what users type

Anything that arrives from outside, a form field, an address, an uploaded file, may contain instructions disguised as data. If the app builds a database question by gluing the user's text into it, a clever visitor can change the question and read or delete things they should not touch. If it displays user text on a page without cleaning it, one visitor can plant a script that runs in other people's browsers.

These are old problems with well-known cures, and assistants know the cures. They simply do not always apply them, especially in the quick, "just make it work" parts of the code. The fix is mostly mechanical: use the safe way of asking the database, clean everything that is displayed, and check every input on the server, not only in the browser.

4. Software packages that do not exist

Modern apps are built from hundreds of ready-made components. When an assistant needs one, it sometimes invents a plausible name for a component that does not exist. Attackers know this. They register the invented names and fill them with harmful code, so that the next person who accepts the suggestion installs it without a second thought. The tactic has acquired a name of its own, and it works because the dangerous package looks exactly like a legitimate one.

The defence is boring and effective: check that every component is real, well established and from its official source; fix its version so it cannot change silently; and run a routine scan for components with known weaknesses. A short list of what your app is made of, kept up to date, is worth more than any tool.

5. Security checks that only look like checks

To get a feature working, an assistant sometimes writes a stand-in: a function called "check permission" that always says yes, a login that accepts any password during testing, a note saying "to do: validate". Most of the time the stand-in is supposed to be replaced. When nobody is reading the code, it is not.

These are easy to find if somebody looks and nearly impossible to find if nobody does. A line-by-line read of the parts that handle logins, payments and personal data, done by a person, catches almost all of them.

6. Doors left open: admin pages and missing limits

Assistants often create convenient extras: an admin page, a test endpoint, a route that lists all users. In development they are useful. In production, if they are reachable and unprotected, they hand out the keys. A similar gap is the missing limit: a login form that lets a program try ten thousand passwords a minute, a contact form that can be used to send unlimited spam, a search box that can be used to overload the server.

Neither problem shows up when you use the app normally. Both appear immediately when somebody who is not being polite tries it. The cure is to list every address the app answers to, decide who may reach each one, and put sensible limits on anything a program could repeat.

7. Hidden instructions for the assistant itself

Teams that work with AI assistants often share small configuration files that tell the assistant how to behave: coding style, preferred libraries, rules. Those files are a target. Instructions can be hidden inside them using invisible characters, so that a file looks harmless to a person but quietly tells the assistant to add a weakness, fetch something from the internet or leave out a check. The same trick can be used through documents or web pages the assistant is asked to read.

You cannot see this problem by eye, which is the point. It is handled by treating those shared files like code: review every change, accept them only from known people, and use tools that reveal hidden characters. It is also a good reason not to let an assistant act on content from the open internet without supervision.

8. Unsafe handling of files and data formats

Apps read and write many kinds of files and data. Some formats are inherently risky: reading a file in the wrong way can let the file run code on your server, and processing images, spreadsheets or archives without limits can crash the app or expose other files on the machine. These flaws live deep in the plumbing, they have no obvious user-facing symptom, and they are the kind of thing an assistant reproduces from examples it has seen.

Allow only the file types you expect, limit their size, never open an uploaded file with something that can execute it, and keep uploads away from the part of the server that runs the app. A specialist review catches the rest.

9. Careless use of personal data

The ninth problem is not a flaw in the code but in the way the work is done. Real customer data is pasted into an assistant "to get a realistic test". A copy of the live database is used for development. Logs record whole requests, including passwords or card details. Backups sit unencrypted in a folder anyone can open.

For a business, this is where a technical slip becomes a legal one. Personal data is protected by law, in the UK and in the European Union alike, and the obligations apply to a prototype as much as to a finished product. The simple rule is that test environments use invented data, and real data goes only where there are real protections.

What it means for your business

If you are not a developer, three consequences are worth keeping in mind.

First, your customers will ask. Larger customers increasingly send security questionnaires before they sign, and they want to know who built the software, how it was reviewed and where the data lives. An app that was never reviewed cannot answer honestly.

Second, a breach is expensive in ways that have nothing to do with code. You may have to tell the data protection authority and the people affected, you will spend time and money investigating, and trust takes much longer to rebuild than software takes to fix. For a small company, the damage to reputation is often larger than the fine.

Third, responsibility does not move to the tool. If an AI wrote the flaw, you are still the one who published it. "The assistant did it" is not a defence either with a customer or with a regulator.

The question to ask any builder"Who reviewed the permissions and the handling of personal data, and how was it tested on the running app?" A good answer names a person and a method. A vague answer is information too.

How to find out whether your app has these problems

You can learn a lot without reading a line of code. Five questions, asked of whoever built the app, separate the prepared from the hopeful:

  • Where are the passwords and keys kept, and have any ever been in the shared code?
  • If I create two ordinary accounts, can either one reach the other's data by changing something in the address?
  • Which pages and addresses exist that a normal user never sees, and who can open them?
  • What is the app built from, and when were those components last checked?
  • Is there any real customer data outside the live system, in tests, copies or logs?

For a proper check, the independent review has to look at the running application as an outsider would. Reading the source finds some problems, but permissions, hidden pages and weak limits are best found by trying: logging in as different people, altering requests, expiring sessions and feeding the forms unexpected input. A finding is only closed when the same test no longer succeeds.

Automated scanners are a useful first pass. They are good at spotting exposed keys and known-vulnerable components, and poor at the problems that depend on how your business works, such as who may see which invoice. That second group needs a person who understands your rules.

Fix it or rebuild it?

Once problems are found, the practical question is whether to patch the app or start the affected part again. A useful rule of thumb comes from asking one question: can all the "who may do what" decisions be described in one place?

  • Repair when the problems are local: a key to replace, a query to correct, a component to update, an unprotected page to close. Each fix is small and independent.
  • Rebuild the permission layer when the checks are scattered across dozens of screens with no pattern. Patching them one by one is slow, and a single missed screen puts you back where you started. A single central rule that every screen consults is safer and cheaper to maintain.
  • Rebuild the data layer when the app serves several customers from one database and was never designed to keep them apart. Separation added afterwards is fragile.
  • Consider starting again when nobody, including the original builder, can explain how the app works. Code that nobody understands cannot be made trustworthy, only hoped about.

The prototype is not wasted in the last case. It is an excellent specification: it shows exactly what people liked and what the screens should do. The new version keeps that knowledge and builds it on a foundation that was designed rather than accumulated.

Keeping it from happening again

The goal is not to stop using AI assistants. It is to put a few cheap habits around them.

  • Write the security rules down before asking for the feature. Who may see this? What may be typed here? What must never be stored? A short list in the request changes the result noticeably.
  • Keep changes small. A change that can be read in ten minutes can be reviewed. A change that cannot, will not be.
  • Treat generated code as input from a stranger. It may be excellent, but it is checked before it is trusted, especially around logins, money and personal data.
  • Automate the boring checks. Secret scanning, component scanning and static analysis can run on every change, and they never get tired.
  • Test the running app with different roles. This is the single most valuable habit, and the one most often skipped.
  • Keep real data out of experiments. Invented data for tests, and real data only in the live system.

How we work with AI assistants

We use AI assistants every day to build software, and we would not want to go back. We also treat the output the way we would treat a quick draft from a very fast junior colleague: useful, tireless, and in need of a second pair of eyes. A person owns every design decision, a person reads what goes near logins, payments and personal data, and the finished application is tested as an outsider would test it. Customer data does not go into our assistants or into test copies.

We do not have a benchmark of our own to quote, and we would rather say so than invent one. What we can say is that the habits above are cheap, they are the ones we apply to our own products, and the problems they prevent are the ones that cost other people the most. If you already have an app built this way and are unsure about it, the first step is an independent look at the running application. We are happy to help with that, and we will tell you plainly what is fine and what is not.

Common questions

Is it a mistake to build an app with AI tools?

No. It is a fast, effective way to build, as long as the result is reviewed and tested before real people rely on it. The mistake is skipping that step because the demo looked good.

Can I just ask the AI to check its own code?

It helps a little, and it is worth doing as a first pass, but it is not enough. The same model tends to repeat its own blind spots. An independent review, with testing of the running app, is what gives real confidence.

Do free scanners solve the problem?

They catch a useful part of it: exposed keys and known-vulnerable components. They are weak at problems that depend on your business rules, for example who may see which record, so they complement a human review and do not replace it.

When is a professional review really necessary?

Before the app handles real customer data or payments, after any large AI-generated change, and always in regulated fields such as health or finance. For a personal tool that touches nothing sensitive, a lighter check is reasonable.

Does this apply to small internal tools?

Yes, in proportion. An internal tool that holds customer names or connects to company accounts deserves the same basic checks: where the keys are, who can open it, and what data it keeps.

Photo: FlyD on Unsplash

S

Softbee Labs. We build software and private AI for small and medium businesses, and write down what we learn along the way. See how we work.

One useful email a month.

New articles and experiments from the Labs, in plain language. You can leave at any time.