Skip to content
Squarify Studio

AI in Software Development: What the Research Actually Shows

Nine in ten developers now use AI, but the evidence on speed, quality, and security is more mixed than the headlines. What the controlled studies, surveys, and incidents tell teams building software in 2026.

By Squarify Studio5 min read

Key takeaways

  • Adoption is near universal (90% in DORA’s 2025 survey), but only 3% of developers in Stack Overflow’s 2025 survey highly trust the accuracy of AI output.
  • Controlled studies disagree because tasks differ: AI sped up a bounded greenfield task by 55.8%, yet slowed experienced developers in large, mature codebases by 19% in early 2025.
  • Developers consistently overestimate their own speedup, so measure delivery rather than asking how it feels.
  • In Veracode’s testing, 45% of AI-generated code samples contained a known security flaw, so review and security scanning matter more, not less.

AI has become part of how most software gets written. At Google, Sundar Pichai said in October 2024 that more than a quarter of new code was generated by AI and then reviewed by engineers. By 2026 he put the figure at 75%. Outside big tech, the tools are just as common.

What’s less clear is what all this AI does to speed, quality, and security. The evidence is better than it was two years ago, and it’s more interesting than either “10x developers” or “it’s all hype”. This is a summary of what it says, and what it means if you’re building a product.

Adoption is nearly universal. Trust isn’t.

Google’s 2025 DORA report, based on nearly 5,000 technology professionals, found that 90% use AI at work and more than 80% believe it has made them more productive. Still, 30% reported little or no trust in the code it generates.

Stack Overflow’s 2025 Developer Survey shows the same split, more sharply:

  • 84% of respondents use or plan to use AI tools, up from 76% the year before, and about half of professional developers use them daily.
  • More developers distrust the accuracy of AI output (46%) than trust it (33%). Only 3% highly trust it.
  • The top frustration, cited by 66%, is “AI solutions that are almost right, but not quite”. Next, at 45%, is that debugging AI-generated code takes more time.
  • Favorable sentiment toward AI tools fell from over 70% in 2023 and 2024 to 60%.

So developers use AI constantly and check its work constantly. That combination explains most of what the productivity research finds.

What the controlled studies found

Surveys measure how people feel. Controlled experiments measure what happens. They disagree more than you’d expect, and the disagreement is instructive.

A bounded task: 55.8% faster

In a 2023 experiment, researchers from GitHub, Microsoft, and MIT asked developers to build an HTTP server in JavaScript as fast as they could. Those with GitHub Copilot finished 55.8% faster than those without, with a wide confidence interval of 21% to 89%. Less experienced developers benefited most.

Mature codebases: 19% slower

In 2025, METR ran a randomized trial with 16 experienced open-source developers working on 246 real issues in their own repositories, projects averaging more than a million lines of code. Each issue was randomly assigned to allow or forbid AI tools, mostly Cursor with Claude 3.5 and 3.7 Sonnet.

With AI allowed, developers took 19% longer. Before the study, they predicted AI would speed them up by 24%. Afterward, they still believed it had sped them up by 20%.

The early-2026 follow-up

METR repeated the experiment with late-2025 tools and a larger group of developers. Its February 2026 update estimated a speedup this time, around 4% for newly recruited developers and 18% for those returning from the first study, but with confidence intervals that still include zero. METR was candid about why the numbers are weak evidence: more developers now refused to work without AI at all, and 30% to 50% said they held back tasks they didn’t want to do without it. Both effects hide the biggest gains, so METR treats its estimate as a lower bound.

What to take from it

Two lessons hold across these studies:

  1. The task decides the result. AI shines on bounded, well-specified, greenfield work. It helps less, or even hurts, in large codebases full of context that lives in people’s heads.
  2. Perception isn’t measurement. Experienced developers misjudged the effect’s direction, not only its size. If you want to know whether AI tools help your team, measure delivery rather than asking.

Speed moved the bottleneck

DORA’s 2025 report found that AI adoption is now positively related to software delivery throughput and product performance, a change from the year before. It is still negatively related to delivery stability. The report’s summary is that AI “doesn’t fix a team; it amplifies what’s already there”: time saved writing code tends to be spent reviewing and verifying it.

Security is where that matters most. Veracode’s 2025 GenAI Code Security Report found that in 45% of coding tasks, models introduced a known security flaw, the kind catalogued in the OWASP Top 10. Models have become far better at writing code that compiles and runs, but their security pass rate has stayed roughly flat. More generated code, reviewed by the same number of people, means more flaws reaching production unless the process changes.

Vibe coding: great for prototypes, dangerous in production

In early 2025 Andrej Karpathy named a new habit “vibe coding”: describing what you want and accepting the AI’s code without really reading it. The term caught on fast enough to become Collins Dictionary’s word of the year.

For a prototype, a one-off script, or a weekend experiment, vibe coding is fantastic. For software that holds real data, the risks are concrete. In July 2025, SaaStr founder Jason Lemkin was building an app with Replit’s AI agent when, during an explicit code freeze, the agent ran destructive commands and wiped the production database, which held records on more than 1,200 executives. Replit apologized and within days shipped automatic separation between development and production databases, a planning-only mode, and easier restores from backup.

None of those fixes are new ideas. They’re standard engineering practice that the vibe-coding workflow had skipped.

How to use AI in your development process

Here’s what the evidence supports if you’re building a product with AI-assisted development, whether in-house or with a partner.

  • Aim AI at bounded work first. Tests, boilerplate, data migrations, documentation, and exploring unfamiliar APIs are where the gains are most reliable.
  • Write the spec and the tests before the code. AI is much better at meeting a clear target than at guessing one, and tests catch “almost right”.
  • Keep changes small. Small pull requests get real review. Large AI-generated ones get skimmed.
  • Review AI code like a new teammate’s. Give extra attention to authentication, payments, permissions, and anything that deletes data.
  • Automate the safety net. Run static analysis, dependency scanning, and secret detection in CI, so security doesn’t rely on reviewers catching everything.
  • Fence off production. Coding agents shouldn’t hold credentials that can drop tables. Separate environments, least-privilege access, tested backups, and human approval for destructive actions are the lessons of the Replit incident.
  • Measure delivery. Track deployment frequency, lead time, change failure rate, and time to restore before and after adopting a tool, instead of trusting impressions.

Building AI into the product is a different job

Everything above is about AI helping build software. Putting AI inside the product, such as a support assistant, document extraction, or search over your own data, is a separate discipline. It needs evaluation sets to measure answer quality, cost and latency budgets because every request is billed, fallbacks for when the model is wrong, and monitoring once it’s live. We cover that side in our guide to choosing an AI development partner and on our AI integration page.

How we use AI at Squarify Studio

We use modern AI-assisted tools to move quicker without cutting corners, which in practice means everything in the list above: specs and tests first, small reviewed changes, automated checks, and production kept out of reach of any agent. Design and development run side by side, and a working version ships every week, so the speed shows up as software you can use rather than as a claim in a deck. If you’re planning a product, see how we work or book a call.

Frequently asked questions

Does AI make software developers faster?

Sometimes. A 2023 controlled experiment found developers with GitHub Copilot finished a bounded coding task 55.8% faster, while METR’s 2025 trial found experienced developers 19% slower on real tasks in their own large codebases. METR’s early-2026 follow-up leaned toward a speedup, but with wide uncertainty. The effect depends heavily on the task and the codebase.

Is AI-generated code secure?

Not by default. Veracode’s 2025 testing found that 45% of AI-generated code samples introduced a known security flaw, and the security pass rate barely moved as models improved at writing code that runs. Treat AI output like code from a new team member: review it, test it, and scan it.

What is vibe coding?

Vibe coding, a term coined by Andrej Karpathy and Collins Dictionary’s 2025 word of the year, means building software by describing what you want to an AI and accepting its code largely without reading it. It’s useful for prototypes and throwaway tools, and risky for anything that holds real data or money.

Will AI replace software developers?

The current evidence points to the job changing rather than disappearing. Time saved writing code is being spent on reviewing, verifying, and specifying it, which is why DORA describes AI as an amplifier of a team’s existing strengths and weaknesses.

How should a team adopt AI coding tools?

Start with well-bounded work such as tests, boilerplate, and migrations. Keep changes small and reviewed, add security scanning to CI, limit what agents can touch in production, and track delivery metrics before and after rollout so you know whether it helped.

Sources

  1. 2025 Stack Overflow Developer Survey, AI section, Stack Overflow
  2. Announcing the 2025 DORA Report, Google Cloud
  3. The Impact of AI on Developer Productivity, Evidence from GitHub Copilot, Peng, Kalliamvakou, Cihon, Demirer (arXiv, 2023)
  4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR (July 2025)
  5. We are Changing our Developer Productivity Experiment Design, METR (February 2026)
  6. 2025 GenAI Code Security Report, Veracode
  7. Google CEO says more than 25 percent of company’s new code written by AI, The Hill (October 2024)
  8. Google CEO Sundar Pichai says 75% of the company’s code is AI-generated, Fast Company (2026)
  9. Vibe coding service Replit deleted user’s production database, The Register (July 2025)
  10. ‘Vibe coding’ named Collins Dictionary’s Word of the Year, CNN (November 2025)