All guides

Workflows · Intermediate

How to Use AI for Programming (Without Shipping Broken Code)

Five concrete developer workflows with real prompts and code — debugging, test generation, code review, learning an unfamiliar codebase, and refactoring — plus the review checklist that matters.

Simanta Pratim DasPublished 9 min read

AI is genuinely useful for programming, but the popular framing — "describe an app and get an app" — is the least reliable way to use it. The workflows that hold up in real projects are narrower and more boring.

The governing principle: AI is strongest where you supply the code and ask for a transformation, weakest where it must invent both the code and the context.

Here are five workflows with actual prompts, plus the review checklist.

Workflow 1: debugging with a hypothesis-first loop

The common approach — paste an error, get a fix — often produces a change that silences the symptom without addressing the cause. Ask for causes before fixes.

The prompt:

This function intermittently returns stale data in production but never
fails locally.

[paste function]

Context: Node 20, single process, ~200 req/s, Redis cache in front of Postgres.
It started after we added the cache layer.

List the 4 most likely causes ranked by probability, given that it's
intermittent and environment-specific. For each: what would confirm or
rule it out. Do NOT propose a fix yet.

Why it works. "Intermittent" plus "only in production" is diagnostic information that narrows the space enormously — it points at concurrency, caching, or load, and away from logic errors. Asking for ranked hypotheses with tests keeps you in control of diagnosis, which is the part requiring judgment. Withholding the fix request prevents the model from anchoring on the first plausible change.

Then, after you have confirmed the cause:

Confirmed: two requests can both miss the cache and write back in a
non-deterministic order. Fix this with a single-flight pattern.
Constraints: no new dependencies, keep the existing function signature,
and explain the failure mode your fix does NOT cover.

That last clause is the highest-value sentence in the prompt. It surfaces the residual risk instead of leaving you with false confidence.

Workflow 2: test generation from real code

This is among the most reliable uses of AI, because the source of truth is in front of it.

The prompt:

Write unit tests for this function using Vitest.

[paste function]

Requirements:
- Cover the happy path, then every early return
- Include edge cases: empty array, single element, duplicate ids,
  null createdAt, and 10k elements for performance sanity
- No mocking of internal logic — only mock the clock
- Name tests as behaviour statements, not "test1"

After the tests, list any branch you could NOT cover and why.

Why it works. The function body is the specification, so nothing needs inventing. Enumerating edge-case categories is what raises coverage — left alone, models write happy-path tests. The final instruction is a coverage audit that regularly reveals untestable code, which is itself a design signal worth acting on.

What to check: that assertions actually assert something meaningful (expect(result).toBeDefined() is theatre), that tests fail when you deliberately break the function, and that mocks do not simply restate the implementation.

Workflow 3: code review as a second reader

Useful precisely because it never gets bored or defers to you socially.

The prompt:

Review this pull request the way an on-call engineer who will be paged
by this service would.

[paste diff]

Prioritise, in this order:
1. Failure modes under load or partial failure
2. Unhandled errors and silent catches
3. Security: injection, authz gaps, secrets in logs
4. Race conditions

Ignore formatting and naming style entirely.
For each finding: severity, the specific line, and why it matters.
If you find nothing in a category, say so explicitly.

Why it works. "Review this code" produces style commentary, because that is the most common form of review text online. The on-call framing implies a value ranking — reliability over elegance — and the explicit priority list plus "ignore formatting" suppresses the low-value noise. Requiring an explicit "nothing found" per category prevents silent skipping.

Its real strengths in review: spotting unhandled promise rejections, missing await, off-by-one errors in pagination, inconsistent error shapes, and forgotten authorisation checks on new endpoints. Its weakness: it cannot know your system's invariants, so it will miss anything requiring architectural context.

Workflow 4: understanding unfamiliar code

The use case that saves the most time and carries the least risk — you are asking about code that exists.

The prompt:

I'm new to this codebase. Explain this module.

[paste file]

Structure your answer:
1. What it's responsible for, in two sentences
2. Data flow: what comes in, what goes out, what it mutates
3. Non-obvious decisions and what problem each probably solves
4. What I'd likely break if I changed it carelessly
5. Three questions I should ask the original author

Mark clearly anything you inferred rather than read.

Why it works. Sections 3 and 5 are the valuable ones. Unfamiliar code is hard not because syntax is confusing but because the reasons are invisible — and a model that has seen enormous amounts of code is good at recognising "this looks like a workaround for a specific class of problem." The final line matters: it separates what the file says from what the model is guessing, which is exactly the boundary you need when you cannot yet tell the difference.

Workflow 5: refactoring with a behaviour contract

Refactoring is risky to delegate because "improve this" invites silent behaviour changes. Fix that by constraining behaviour first.

Refactor this function to reduce nesting.

[paste function]

Hard constraints:
- Identical observable behaviour, including for invalid input
- Same function signature and same thrown error types
- No new dependencies
- Do not "fix" bugs you notice — list them separately instead

Output: the refactored function, then a table of every behavioural
difference you're uncertain about.

Why it works. "Do not fix bugs you notice" is the critical line. Without it, models silently correct perceived bugs while refactoring, which is how a supposedly behaviour-preserving change breaks production. Separating the bug list keeps the refactor reviewable as one concern.

Then verify properly: run the existing tests before and after, diff the output on real input, and read the diff yourself. Never accept a refactor on the basis that the explanation sounded right.

The review checklist

Before any generated code merges:

Existence checks

  • Every function, method, and parameter exists in the official docs
  • Every imported package exists in the registry — check before installing
  • The API version matches what you actually run

Correctness

  • You have run it, including edge cases
  • Error paths are handled, not silently swallowed
  • Boundaries are right: empty, null, one element, very large

Security

  • Inputs validated at the trust boundary
  • Queries parameterised, not interpolated
  • No secrets in logs or error messages
  • Authorisation checked on new endpoints, not just authentication

Comprehension

  • You can explain every line to a colleague
  • You know what happens when it fails

That last pair is the real gate. Code you cannot explain is code you cannot maintain or debug at 3am, no matter how well it works today.

Two specific hazards

Fabricated packages. Models invent plausible package names, and attackers have registered names that models commonly hallucinate. Verify in the registry before npm install. This is an active supply-chain attack vector, not a theoretical one.

"I tested this." No test was run. No code was executed. Treat such statements as generated text.

What not to delegate

  • Architecture decisions. These depend on team, scale, deadlines, and existing systems — context the model does not have. Use it to enumerate trade-offs, then decide yourself.
  • Cryptography and auth flows. Subtly wrong implementations look correct. Use audited libraries.
  • Anything touching money, permissions, or personal data without line-by-line review.
  • Code you will not read. The fastest way to accumulate unmaintainable systems.

Using this with PhantomAI

PhantomAI supports these workflows through the parts that matter here: you can paste or attach code and ask questions against it, syntax-highlighted output for readability, and a canvas view for working through longer files. openai/gpt-oss-120b handles the reasoning-heavier work above — hypothesis ranking, review, refactor analysis — while openai/gpt-oss-20b is quicker for small syntax questions.

Be clear on one thing: PhantomAI does not execute your code. It cannot run tests, reproduce a bug, or verify that a suggested fix works. Every claim about behaviour is a prediction you must confirm by running it. See our limitations page for the full list.

Continue

Simanta Pratim Das

Founder & Developer

Simanta is an independent AI engineer based in Guwahati, India, building PhantomAI as a solo project — designing the product, the interface, and the AI pipeline end to end.

Related guides

Try these techniques in PhantomAI

PhantomAI is free to use, and the workflows in this guide work best when you paste in your own material rather than relying on the model's memory. Before you start, it's worth reading what it can't do.