AI Can Find the Bug. Your Career Signal Is Whether You Can Close It.

AI can now hand you a long list of suspicious code before your first coffee. That does not make you a strong reviewer. The stronger career signal is what happens after the alert: whether you can reproduce the failure, establish who and what it affects, separate urgency from theater, give the right owner usable evidence, and verify that the fix closes the path without creating another one.
This distinction became unusually visible around the Linux 7.2 release candidates. In his August 9 announcement, Linus Torvalds described a large late-cycle release candidate as part of a “new normal” with many fixes resulting from review by various AI tools.[1] That is a meaningful capability signal, but it sits beside an operational problem: automated analysis can also produce duplicates, weak reports, and more triage work than maintainers can absorb.
The Linux kernel's own security guidance makes the receiving side concrete. A report needs technical detail that helps developers diagnose the issue, and AI-found vulnerabilities follow explicit reporting rules rather than automatically entering the private security channel.[2] Finding something is only the opening move. A useful contribution has to fit a real maintenance process.
For job seekers, “I ran an AI reviewer and found 37 bugs” leaves the important questions unanswered. A more defensible way to present the work is to show the downstream judgment required to turn one plausible finding into one responsibly closed problem.
More detection creates a new bottleneck
Traditional code review has always mixed several kinds of work: understanding intent, checking behavior, spotting risk, negotiating tradeoffs, and deciding whether a change is ready. AI review can widen the first-pass search. It can inspect more paths, propose edge cases, and point reviewers toward code they might otherwise miss.
That does not mean every proposed issue is real or important. GitHub described an instructive failure while migrating Copilot code review to a set of better-maintained exploration tools. The change looked like an upgrade, yet internal benchmarks showed higher cost and fewer caught issues until the team studied and reshaped the agent's workflow.[3] Better components did not automatically produce a better reviewer.
GitHub's work on agentic validation offers a useful analogy. In a study of a computer-use agent running extension tests, the agent's self-assessment could not reliably identify “not a bug” scenarios, while an independent structural approach performed better at distinguishing product failures from agent execution errors.[4] This was execution validation rather than code review, so its metrics do not describe AI reviewer accuracy. It does illustrate why independent evidence can be useful before a team accepts an agent's account of what happened.
Detection capacity can therefore move the bottleneck into attention. Every plausible issue asks somebody to reconstruct context, check reachability, understand permissions, search for duplicates, estimate impact, and decide whether to interrupt planned work. If the report lacks evidence, the maintainer pays that cost from scratch.
This is why alert quality becomes a team performance issue. GitHub's secret-scanning work argues that noisy alerts erode trust and slow remediation, while context-aware verification helps distinguish genuine exposures from values that only look sensitive.[5] A reviewer who reduces uncertainty creates capacity. A reviewer who forwards every model suspicion transfers the cost.
The career opportunity is to prove that you can manage that boundary.
The finding-to-fix chain
A credible AI-assisted review story has six stages: detect, reproduce, bound, prioritize, route, and close. Each stage answers a different question, and skipping one usually pushes uncertainty onto somebody else.
The professional contribution is the full evidence chain, not the first alert.
1. Detect: preserve the claim before improving it
Start by recording what the tool actually alleged. Keep the relevant file, line, commit, configuration, input conditions, and explanation. Do not silently rewrite a vague alert into a stronger claim and later attribute the whole analysis to the tool.
This matters for both honesty and debugging. The original finding may contain the right location but the wrong mechanism. It may identify a dangerous function without proving that untrusted input can reach it. It may reproduce on the default branch but not on the tagged version a customer runs.
Preserving the raw claim lets you show the human work that followed. It also helps you evaluate the reviewer over time: which classes of findings were useful, which were duplicates, and which consistently misunderstood the codebase?
A compact detection record might contain:
Finding: tenant ID from request may not constrain export lookup
Source: AI-assisted review of commit 8f3c2a1
Location: exportService.ts, getExport()
Claimed impact: cross-tenant document access
Unknowns: route authorization, query filters, deployed versions
The unknowns are not an embarrassment. They define the investigation.
2. Reproduce: turn suspicion into an observable failure
A report becomes much more useful when another person can observe the same behavior. Build the smallest safe reproducer that tests the claimed path.
For a functional bug, that may be a failing unit test, an integration test, or a minimal input and expected output. For a security issue, work in an authorized environment and avoid touching real user data. Use a local fixture, a test tenant, a mock service, or a deliberately vulnerable sample. If you do not have permission to test a production system, stop before experimentation becomes intrusion.
Reproduction does more than confirm the model. It often changes the diagnosis. A suspicious query might be protected by middleware. A race condition may require a timing window that normal deployment cannot reach. A parser failure may occur only under an unsupported configuration.
Strong evidence records the result clearly:
- Confirmed: the minimal test fails under stated conditions.
- Not reproduced: the proposed path is blocked, and the blocking control is named.
- Inconclusive: the environment or evidence cannot safely answer the question.
“Inconclusive” is a professional result when it prevents an unsupported claim.
3. Bound: identify where the issue starts and stops
Once a failure exists, establish its boundary. Which versions are affected? Which inputs trigger it? Which permissions are required? Does the result expose data, corrupt state, reduce availability, or merely produce a confusing message? Is the path on by default?
Repository context is decisive here. GitHub Security Lab's vulnerability-triage workflow checks whether reports contain consistent evidence and then incorporates repository-specific controls such as permission checks and sanitizers. Those controls can turn an apparent vulnerability into a false positive.[6]
Bounding also prevents exaggerated language. “Authorization bypass” sounds severe. “A user with export-admin permission can request an expired export ID from another tenant in version 2.4 when legacy routing is enabled” is narrower and far more actionable.
If you are writing a portfolio case study, the boundary is often the strongest proof of judgment. It shows that you investigated what would have falsified the alarming interpretation instead of searching only for confirmation.
4. Prioritize: separate severity from novelty
Novelty is exciting. Priority depends on impact, likelihood, exposure, existing controls, and the cost of delay.
An AI reviewer may use urgent language because it recognizes a pattern associated with severe incidents. Your job is to connect the pattern to this system. Ask:
- What outcome is possible?
- Who can trigger it?
- What access do they already need?
- How repeatable is the path?
- Which monitoring or compensating controls exist?
- What could a rushed fix break?
This is not an invitation to minimize risk. It is how teams direct scarce attention toward the most consequential work. A low-complexity data exposure affecting default configurations should not compete equally with a cosmetic exception message in an optional admin tool.
GitHub's CodeQL updates illustrate that detection quality is continuously refined through framework-specific modeling and false-positive reduction.[7] Mature review practice treats classification as engineering work, not clerical cleanup after the “real” discovery.
5. Route: make the receiving team's next step cheaper
A useful report gives the owner the evidence needed to reproduce, judge, and act. Extra length earns its place only when it reduces uncertainty.
Before filing, read the project's reporting policy. Search open and recently closed issues. Check whether the affected branch is supported. Use the requested channel and disclosure path. A private security inbox, public issue tracker, internal incident system, and pull-request comment serve different risk models.
An evidence-rich packet usually includes:
- a one-sentence observed behavior;
- affected version or commit;
- prerequisites and permissions;
- minimal reproduction steps or a failing test;
- expected versus actual behavior;
- impact stated within the demonstrated boundary;
- relevant logs, stack traces, or code path;
- duplicate search performed;
- and a proposed next step, if you have enough context.
Avoid pasting an unedited model transcript. The maintainer should not have to extract the claim from twenty speculative branches. Keep the original analysis in your private record and submit the smallest defensible case.
Good triage compresses uncertainty before it reaches the maintainer.
6. Close: prove the risk changed
An accepted issue is not the end of the story. Closure means the relevant behavior changed and the change is protected against regression.
If you own the fix, trace the cause rather than patching only the sample input. Add a regression test that fails before the change and passes after it. Check adjacent paths that share the same assumption. Review whether the fix changes compatibility, performance, permissions, or user-visible behavior.
If another team owns the fix, contribute the reproducer, answer questions, retest the proposed patch, and update the record. A candidate does not need to author every line to demonstrate ownership. Coordinating a safe closure across boundaries is often the more realistic senior-level contribution.
OpenAI's engineering-team guide describes AI review as a baseline that can catch meaningful bugs while keeping engineers responsible for whether code is ready to ship.[8] That final responsibility is the career signal. The model can suggest closure. The team has to establish it.
A worked example: the dramatic authorization alert
Imagine an AI reviewer comments on a pull request for a document-export service:
Critical:
getExport(exportId)does not filter bytenantId, allowing any authenticated user to download another customer's export.
The comment deserves attention. It does not yet prove the claim.
The weak response is to copy the text into a critical incident, tag security leadership, and report that the AI found a cross-tenant vulnerability. The opposite weak response is to dismiss it because the reviewer has produced false positives before.
The finding-to-fix response starts with the route. You trace the request through authentication middleware and find that normal users receive a tenant-scoped database client. That weakens the original claim. You also discover a legacy admin route that constructs the service with an unrestricted client. That route checks an export-admin permission but does not compare the requested export's tenant with the administrator's assigned tenants.
You create two test tenants and one test administrator authorized only for tenant A. In the test environment, the administrator can request tenant B's export by ID through the legacy route. The ordinary user route remains correctly scoped.
Now the boundary is clear:
- affected surface: legacy admin export route;
- required access: authenticated account with
export-admin; - affected data: completed export documents when the identifier is known;
- unaffected surface: standard tenant user route;
- confirmed version: current staging commit;
- unknown: whether identifiers have appeared in logs or support records.
The report can be prioritized without hype. The required permission reduces exposure but does not make cross-tenant access acceptable. The team disables the legacy route, audits recent access logs, adds an explicit tenant predicate in the service layer, and writes regression tests for both user and admin paths.
The closure record states what changed and what remains uncertain. It does not claim “zero impact” merely because the logs showed no known misuse. It says the team found no matching access in the available retention window.
That is a powerful interview story because the AI alert was neither perfectly right nor useless. It pointed toward a real flaw while misstating the reachable population. Human investigation made the finding precise enough to fix.
Build a portfolio artifact around closure, not surveillance
You do not need access to a company's private code to demonstrate this skill. In fact, indiscriminately scanning public projects and flooding maintainers with machine-generated reports is a poor portfolio strategy.
Use one of three safer routes:
Your own project. Add a deliberately flawed branch or investigate a genuine defect from your issue history. Preserve the review output, build the reproducer, fix the cause, and document the regression test.
An authorized lab. Use a security training application, benchmark, capture-the-flag environment, or sample repository that explicitly permits testing. Focus the case study on method rather than exploit spectacle.
A project where you already contribute. Read its reporting guidance first, coordinate with maintainers, and report only findings you have reproduced within the project's rules. One high-quality contribution is stronger evidence than a dashboard of unverified alerts.
Create a compact closure packet:
evidence/
original-finding.md
reproduction.md
boundary-and-severity.md
fix-decision.md
regression-test.md
closure-note.md
The packet should show where the machine stopped and your judgment began. Note what the reviewer got right, what it got wrong, which evidence changed your assessment, and how the final fix was verified.
This complements a handoff-ready AI workflow. The handoff article asks whether somebody else can operate the system. This closure packet asks whether somebody else can audit one of its most consequential claims.
How to write this work on a resume
The phrase “used AI code review” says almost nothing about your contribution. A strong bullet names the problem, the judgment you supplied, and the changed engineering outcome.
Weak:
Used AI tools to find 37 bugs and improve code quality.
Stronger:
Built an AI-assisted review workflow with reproducibility and duplicate checks, converting high-confidence findings into regression-tested fixes while routing inconclusive cases for owner review.
Weak:
Discovered a critical security vulnerability with an LLM.
Stronger:
Reproduced and bounded an authorization flaw flagged during AI-assisted review, traced exposure to a legacy admin path, and added tenant-scoped regression coverage after coordinating remediation.
Weak:
Automated vulnerability triage with agents.
Stronger:
Designed a repository-aware triage pipeline that enriched scanner alerts with reachability, permission, and sanitizer context, reducing unsupported escalations and giving maintainers reproducible issue packets.
Use numbers only when the measurement is defensible. “Reduced false-positive review time by 28% across 120 alerts” is useful if you defined the baseline and tracked reviewer effort. “Improved security by 80%” is not.
Choose the evidence that matches the role. A security role may emphasize threat modeling and disclosure. A platform role may emphasize alert routing and service ownership. A staff engineering role may emphasize prioritization, cross-team closure, and the regression strategy. A developer-tools role may emphasize reviewer evaluation and feedback loops.
If you need distinct resume versions for those roles, CoreCV's role-targeted workflow can generate a role-focused version and fine-tune it against a job description or job URL. Keep the underlying incident facts fixed. Change which relevant evidence receives space.
The same discipline helps make quiet infrastructure work legible: describe the risk reduced, the operating decision, and the proof that the system changed.
One evidence base should support every career format without changing the facts.
Tell the interview story as an investigation
Do not make the model the protagonist. Open with the uncertainty.
An AI reviewer flagged a possible cross-tenant export vulnerability. The standard route was protected, so the original claim looked overstated. I kept tracing the service construction and found a legacy admin path using an unrestricted client. I reproduced access between two test tenants, bounded it to administrators with a specific permission, and worked with the service owner to disable the route, add tenant checks, review available logs, and create regression coverage. The alert found the location. My contribution was turning it into a defensible scope and a verified closure.
That story invites useful questions:
- What evidence would have made you dismiss the finding?
- How did you avoid exposing real data during reproduction?
- Why was the issue prioritized at that level?
- Which adjacent paths did you test?
- What did the AI reviewer misunderstand?
- How did you know the patch addressed the cause?
- What would you automate next, and what would remain independent?
Good answers show a changing mental model. You began with a broad claim, found a control that narrowed it, discovered an exception that confirmed a smaller flaw, and adjusted the response to match. That is more credible than pretending you knew the answer at first glance.
It also demonstrates the tradeoff reasoning that strong technical interviews reward. As we have argued before, the best candidates explain rejected options and constraints, not only wins.
A seven-day finding-to-fix exercise
You can build this evidence around an existing project in one week.
Day 1: Choose one bounded surface. Pick authentication, input validation, state transitions, error handling, data export, or another area where expected behavior can be stated clearly. Do not scan an unbounded system.
Day 2: Run two review methods. Use an AI reviewer plus a deterministic tool, tests, or manual checklist. Save exact findings without merging them into one narrative.
Day 3: Classify the queue. Mark each item confirmed, not reproduced, duplicate, expected behavior, or inconclusive. Record the evidence for the classification.
Day 4: Reproduce one meaningful finding. Build a minimal test in an authorized environment. Write expected and actual behavior before editing the implementation.
Day 5: Bound and fix. Identify affected versions, inputs, users, and adjacent paths. Address the cause and add regression coverage.
Day 6: Ask another engineer to review the packet. Can they rerun the failure and understand your priority decision without the chat transcript? Revise gaps they identify.
Day 7: Package the career evidence. Write a short case study, one resume bullet, and a two-minute interview account from the same closure record. Include limitations.
The result may be only one fixed bug. That is enough. Depth of closure is the point.
What this signal proves, and what it does not
A careful closure packet demonstrates investigation, technical communication, prioritization, and respect for team process. It does not automatically prove deep security expertise, production incident leadership, or mastery of every code path. Claim only the scope you owned.
AI-assisted review also varies widely by language, repository, model, configuration, and task. GitHub's experience with changed tools is a reminder that reviewer performance must be measured in its actual workflow.[3] Do not transfer one vendor's benchmark or one successful project into a universal promise.
The need for expertise remains. OpenAI's field report on agentic scientific computing found substantial acceleration while emphasizing that expert guidance, understanding, taste, and care were still required for ambitious work.[9] Code review may present a similar pattern: faster exploration can increase the value of the person who knows what evidence should change a decision.
The alert is a lead. The closure is the work.
AI will keep making potential bugs easier to surface. That is good news when findings become fixes, and expensive noise when nobody owns the path between them.
Build your career evidence around that path. Preserve the original claim. Reproduce it safely. Establish the boundary. Prioritize without drama. Route a compact evidence packet. Verify the fix and protect it with regression coverage.
The machine may deserve credit for pointing at the suspicious line. Your professional value is making the result dependable enough for a team to act on.
For one practical signal each week, follow the AI Career Signals archive. The series focuses on what changed, what the evidence supports, and what technical candidates can do next.