AI is changing software development by compressing work that once took days into hours. Code can be generated, tested, refactored, and documented faster, allowing development teams to deliver functionality at a pace that was previously difficult to achieve.
That acceleration creates an important challenge for security teams.
AI-powered software development can accelerate innovation, but it can also introduce hidden vulnerabilities when security implications are not clearly understood. Watch this webinar series on AI threat realities and proactive security readiness to gain firsthand knowledge of what makes this AI threat unique, coupled with remediation guidance for long-term resilience:
This series also explores how organizations can securely adopt and implement AI at business speed without losing control of cyber risk.
The risk is not simply that AI-generated code can contain vulnerabilities. Human-written code has always contained vulnerabilities. The more consequential issue is that AI can generate technically plausible, well-structured implementations without understanding the business assumptions that determine whether those implementations are secure.
A recent penetration test for a financial services organization illustrates the distinction.
The assessment covered a customer onboarding application processing highly sensitive personal and financial information. Portions of the application had been built or heavily accelerated using an AI coding assistant.
The application worked as intended.
The trust model behind part of it did not.
When an Identifier Becomes an Authentication Mechanism
The application needed to solve a legitimate business problem. Prospective customers could begin an application before establishing full account credentials. They could leave partway through the process and return later, so the application required a mechanism to restore their progress and reconnect them with the correct application record.
The implementation used a temporary applicant-access model.
The critical weakness was in how that access was granted.
The application treated possession of an applicant's globally unique identifier, or GUID, as sufficient context to obtain or restore an access token for that applicant.
A GUID can identify a record. It does not establish that the person presenting it is entitled to access that record.
It does not prove control of an authenticated session, trusted device, verified identity, or validated communication channel.
The system therefore blurred two fundamentally different security concepts: identification and proof of entitlement.
If another applicant's GUID could be obtained, the application could issue or restore access without sufficiently establishing that the requester was entitled to act as that applicant. The resulting exposure potentially included names, contact information, application status, financial information, identity verification data, payment information, Social Security numbers, and co-applicant information.
What made the weakness particularly instructive was that the implementation did not appear careless.
The AI-generated implementation plan included temporary tokens, expiration, rate limiting, logging, session restoration, and suspicious-activity detection. These are all legitimate security controls.
But they protected the wrong part of the trust decision.
The question that mattered came earlier: What must a requester prove before the backend issues or restores an applicant access token?
The implementation focused heavily on what happened after token issuance. The vulnerability existed before issuance, when the application decided who was entitled to receive the token.
A short-lived token is still dangerous when issued to the wrong person. Cryptographic signing does not help if the signing service authorizes the wrong requester. Rate limiting may constrain exploitation, but it does not establish authorization. Logging provides evidence after an event, but it does not make the underlying trust decision correct.
The application had security mechanisms. What it lacked was the right security invariant.
AI Can Produce Secure-Looking Code Without Understanding What Must Be True
This distinction has significant implications for AI-assisted development.
Coding assistants are highly effective at reproducing established implementation patterns. Given a request for temporary applicant access, a model can generate routes, middleware, token services, expiration logic, storage mechanisms, logging, tests, and error handling.
The resulting code may be clean and technically sound.
But security depends on more than whether individual components follow secure coding patterns.
It depends on whether the assumptions connecting those components are correct.
In this case, the relevant questions were not simply whether the token expired or whether the endpoint validated its input.
They were:
- What establishes applicant ownership?
- Can an applicant identifier appear in a URL, browser history, application response, support ticket, email, log, or analytics platform?
- Is the identifier considered public, sensitive, or secret?
- What independent evidence proves that the person presenting it controls the applicant journey?
- What information becomes available once the backend accepts that evidence?
- What should happen when a requester presents a completely valid identifier that belongs to somebody else?
Those questions require an understanding of the application beyond its syntax.
They require knowledge of its trust boundaries, customer journey, identity model, data sensitivity, architecture, expected attacker behavior, and business consequences.
An AI model can reason about those issues when given sufficient context. It cannot be assumed to possess that context simply because it has access to the code.
That is where human judgment remains critical.
Business-Logic Security Lives Between the Lines of Code
Traditional application security testing is often strongest when a vulnerability has a recognizable technical signature.
SQL injection, unsafe deserialization, exposed secrets, vulnerable dependencies, and certain insecure data flows can often be detected because something identifiable is wrong with the implementation.
Business-logic vulnerabilities are different.
Individual pieces of the implementation can be perfectly valid.
The endpoint can function correctly.
The token can be properly signed.
The expiration can be enforced.
The database query can be parameterized.
The middleware can execute exactly as designed.
The vulnerability can still exist because the application has made the wrong decision about who is allowed to initiate that sequence.
This is why AI-assisted software development increases the importance of threat modeling rather than reducing it.
As the cost of producing functional code falls, the security bottleneck increasingly shifts from writing code correctly to determining whether the system is making the right decisions.
For security practitioners, this means reviewing not only the implementation but also the assumptions that connect identities, objects, actions, and trust.
For security leaders, it means recognizing that higher development velocity can produce more security-sensitive decisions per unit of time, even when code quality remains high.
AI Can Also Accelerate the Security Review
The same capability that accelerates development can also increase the reach of security teams.
During the penetration test, an LLM was used to assist with reviewing portions of the application's code. The model helped surface security-relevant patterns, trust-boundary issues, and business-logic weaknesses.
Among the issues identified was the applicant-token flow. The analysis recognized that possession of the applicant GUID was being treated as sufficient context for token issuance without stronger proof that the requester controlled the relevant applicant identity, session, device, or communication channel.
This demonstrates an important shift in application security.
AI can lower the cost of first-pass security analysis in much the same way that it lowers the cost of first-pass software development.
A model can examine an unfamiliar codebase and ask questions such as:
- Can this token be issued without proving ownership?
- Is an identifier being treated as if it were a secret?
- Can one user's input affect another user's authorization context?
- Which endpoints cross trust boundaries?
- What sensitive data becomes reachable after this authorization decision?
- Which negative tests are missing?
- What assumptions does this workflow make about identity or session ownership?
These questions can help security teams direct scarce expert attention toward areas that warrant deeper investigation.
The value is not that the model provides a final security verdict. Rather it is that it can accelerate the path to the right questions.
AI-Assisted Review Does Not Remove the Need for Security Expertise
Finding suspicious code and determining business risk are different activities.
An AI-assisted review can surface dozens of possible authorization weaknesses. A security practitioner still has to establish whether they are exploitable and whether they matter.
That requires understanding whether an endpoint is externally reachable, how an identifier can be obtained, which compensating controls exist, what privileges are required, what data or actions become accessible, and what exploitation means to the organization.
Consider three hypothetical authorization findings.
- One exposes an internal identifier with no meaningful security consequence.
- Another allows a user to perform an action that appears privileged in the code but has deliberately been delegated by the business.
- A third allows an unauthenticated requester to retrieve another customer's regulated financial information.
A model may initially classify all three as authorization concerns but it cannot determine the consequences of each finding. That is the role of human judgment.
It requires connecting the technical finding to the asset, identity, business process, attack path, regulatory exposure, and potential operational impact.
AI can accelerate analysis. It cannot be given responsibility for business context that has never been made explicit.
Secure AI-Assisted Development Starts With Security Invariants
Organizations adopting AI coding tools should therefore avoid treating secure AI development as simply another scanning problem.
The security model needs to be defined before implementation.
Consider a prompt that asks an AI assistant to:
"Build a mechanism that allows applicants to resume an incomplete application."
The requirement describes functionality, but it leaves the model significant freedom to determine how trust should work.
A stronger specification would establish the security invariants alongside the functional requirements:
- Applicant identifiers must never constitute proof of identity or ownership.
- Token issuance must require independent proof of applicant control.
- Temporary authorization must be scoped to the verified applicant.
- Session restoration must not succeed based only on possession of a GUID or other record identifier.
- One applicant must never be able to obtain access associated with another applicant.
- Negative tests must verify cross-user isolation, not only successful restoration.
- Sensitive data must not become accessible until the required level of identity or session assurance has been established.
In the second scenario, AI is not being asked to invent the security model while implementing the feature. It is being asked to implement a feature within explicit trust constraints.
This principle extends beyond onboarding applications.
Authentication, account recovery, authorization, payments, administrative actions, document access, data exports, delegated access, and other sensitive workflows all depend on business-specific security invariants.
Those invariants should be established by people who understand both the architecture and the consequences of failure.
Code Review Needs to Move Up a Level
AI-assisted development also changes what human code review should prioritize.
Traditional reviews often concentrate on implementation details such as unsafe functions, injection risks, insecure cryptography, dependency weaknesses, secret handling, and coding errors. Those checks remain essential.
But when AI can generate large amounts of plausible code quickly, reviewers also need to examine the decisions embodied by that code.
For authentication and authorization workflows, several questions become particularly useful:
What is being proven?
What evidence must a requester provide before receiving access?
- What is merely being identified?
- Are IDs, email addresses, account numbers, GUIDs, or other lookup values accidentally being treated as credentials?
- Where is the trust boundary?
- At what exact point does untrusted input become an authenticated or authorized identity?
- What happens with valid but unauthorized data?
- Does the system reject a legitimate identifier when it belongs to somebody else?
- What capability is unlocked?
- What data, transactions, documents, administrative actions, or downstream services become accessible after the decision?
- What negative tests exist?
- Does the test suite demonstrate that prohibited relationships fail, or does it only demonstrate that expected workflows succeed?
- What is the business impact if the assumption is wrong?
- Would failure expose low-value metadata, regulated PII, payment capability, administrative authority, or another critical business asset?
These questions ask about whether the system's decisions are safe, not whether the code compiles.
Negative Testing Becomes More Important at AI Speed
AI-generated implementations naturally optimize toward the requested success case.
If the prompt asks for application restoration, the implementation and generated tests are likely to demonstrate that restoration works.
Security teams need to deliberately test the opposite.
Can Applicant A restore Applicant B's application?
Can a valid identifier be replayed from a different session?
Can an expired or previously used continuation mechanism be reused?
Can one identity influence the authorization context of another?
Can a request bypass the intended proof-of-control step by calling a lower-level endpoint directly?
Can information disclosed earlier in the workflow later become an authentication factor?
These tests are particularly valuable because they examine the security assumptions between components rather than individual functions in isolation.
AI can help generate these tests.
Humans still need to determine which negative conditions represent meaningful security boundaries for the business.
Security Governance Has to Match Development Velocity
For security leaders, the broader issue is one of scale.
AI allows development organizations to produce more code, make more changes, and introduce more application behavior in the same period of time.
Security programs cannot respond simply by asking human reviewers to inspect proportionally more lines of code.
The review model itself has to change.
AI can be used to expand first-pass analysis, identify sensitive code paths, generate abuse cases, examine authorization relationships, and suggest negative tests.
Human expertise can then be concentrated where context matters most.
Sensitive workflows should receive greater scrutiny, particularly those involving authentication, authorization, account recovery, session restoration, payments, administrative operations, document access, data exports, and regulated information.
AI-generated or substantially AI-modified changes in these areas should be evaluated against explicit security requirements and trust assumptions, not simply coding standards.
The objective is not to slow development back to pre-AI speeds.
It is to ensure that security reasoning scales alongside development.
The Security Bottleneck Is Moving From Execution to Judgment
AI is removing operational bottlenecks throughout software engineering.
Code generation is faster. Testing is faster. Documentation is faster. Refactoring is faster. Security analysis can be faster as well.
The result is not that human expertise becomes less important.
Its value shifts.
When execution is expensive, much of the work is spent producing and inspecting artifacts manually.
When execution becomes cheap, the difficult questions become:
- What should the system trust?
- What must always be true?
- Which failure modes matter?
- Which findings deserve attention first?
- What is the actual business consequence?
- Where should automation be allowed to make a decision, and where is human validation required?
These are judgment questions that require technical expertise, adversarial thinking, and business context.
AI Speed Requires Human Context
The most useful lesson from the penetration test is not that AI writes insecure code.
That conclusion would be too simplistic.
AI can produce secure code, insecure code, and code that is technically sophisticated but based on an incorrect trust assumption. Human developers can do the same.
What changes with AI is the scale and speed at which those decisions can be implemented.
The financial services application had tokens, expiration, logging, and rate limiting. The individual controls looked reasonable.
The failure existed one level above them.
The system did not adequately answer: Why should this requester be trusted with this applicant's access?
No downstream security mechanism could compensate for getting that decision wrong.
AI should be used to accelerate software development. It should also be used to accelerate threat modeling, code review, negative testing, and vulnerability analysis.
But the faster operational processes become, the more important it is to preserve human judgment at the points where technical decisions become business risk.
AI can generate the implementation.
AI can analyze the implementation.
AI can surface the suspicious trust relationship.
Security practitioners still need to determine whether the trust model is valid, and security leaders still need to ensure that the organization has the processes, expertise, and context required to make that judgment at AI speed.
About the author: Zach Mead is a Penetration Tester Expert at Sygnia, where he specializes in advanced web application penetration testing and secure code review. His experience spans both defensive and offensive security, including incident response, application defense, network penetration testing, and OT security assessments. He has conducted complex assessments for organizations ranging from startups to global enterprises and helped secure critical infrastructure in the utilities and maritime sectors. Prior to joining Sygnia, he served as Founder and Principal Consultant at Harbor’s Edge Consulting LLC. At Sygnia, he focuses on emulating sophisticated cyber adversaries to uncover critical vulnerabilities and strengthen defenses across complex environments.
Zach Mead — Penetration Tester Expert at Sygnia https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEjhvp48AIkdttANmLRpuxcKXHlVf_NrvjF43jBH1wmHbB0UtB0tt4yaMrdovhCEXbgAMgpZSsSZDq5UCQiikNuy9VkwuoQN3-geHO2KxXtv7kBMpYI6-8nFwEW0kwtZH4H6ypGfWzEg0TpUnEVf7t0I4TDgmdwdUDNYr86OHTjoM7EdqN6llyAuNJTVJy0/s1600/Zach.png


