← All guides

AI Coding Agent Ecosystem

The tools from prompt to production, layer by layer, with the interview questions each layer attracts and how to answer them with trade-offs.

Source: Infographic by Rathnakumar Udayakumar (@rathanuday). Tool landscape as of 2026-10; check current versions before an interview.

How to answer anything here

  1. Claim: state your pick in one sentence.
  2. Trade-off: name what you give up and the alternative you rejected.
  3. Evidence: say how you would measure it (eval set, latency, cost per task).
  4. Fallback: say what you do when it fails or the vendor changes.

1. Foundation Model

The reasoning engine. Interviews test whether you choose by measured fit, not by brand.

  • Claude
  • GPT-5
  • Gemini
  • Llama
  • Mistral
  • DeepSeek
How do you choose a model for a coding agent?
What they are testing
Do you evaluate on your own tasks, and think in cost per finished task rather than price per token?
A strong answer
List the criteria: task fit, tool-calling reliability, context window, latency, cost per completed task, licence and data rules. Build a small eval set from real tasks, run candidates on it, then route: a cheap model for simple steps, a stronger one for hard steps.
Red flag
"Use the newest or biggest model." No eval, no cost view.
Open-weight or hosted API?
What they are testing
Can you reason about control versus operational burden?
A strong answer
Hosted wins on speed to ship and top quality. Open-weight (Llama, Mistral, DeepSeek) wins when you need data control, fixed cost at volume, or fine-tuning, but you own serving, GPUs and upgrades. Many teams start hosted and move a stable, high-volume step to open-weight once evals prove parity.
Red flag
Treating one option as always better, or ignoring who runs the servers.
A provider deprecates your model next month. What happens?
What they are testing
Vendor lock-in and change management.
A strong answer
Calls go through one gateway layer so the model is config, not code. Pin versions, keep an eval suite, run the replacement against it, shadow-test a slice of traffic, then switch with a rollback path.
Red flag
Provider SDK calls scattered through feature code.

2. AI Coding Assistant

Day-to-day tools. Interviewers want judgement about trusting generated code.

  • Claude Code
  • GitHub Copilot
  • Cursor
  • Windsurf
  • Cline
  • Aider
How do you keep AI-generated code trustworthy?
What they are testing
Engineering discipline, not tool enthusiasm.
A strong answer
Small diffs, tests written or reviewed first, type-check and lint in CI, human review of every change, and a rules file that tells the assistant your conventions. Treat output as a junior contribution: verify, never assume.
Red flag
"I just accept it if it runs."
Compare an inline assistant, an IDE agent and a terminal agent.
What they are testing
Do you match the tool to the task?
A strong answer
Inline completion suits boilerplate while you type. An IDE agent suits multi-file edits you want to watch. A terminal agent suits larger tasks that run commands and tests. Pick by task size and how much oversight you need.
Red flag
Listing product names with no difference in how you would use them.
Tell me about a time an AI assistant was wrong.
What they are testing
Honesty and learning, a behavioural question in disguise.
A strong answer
Use STAR. Say what it got wrong, how you caught it (test, review, reading the diff), the fix, and the guardrail you added afterwards.
Red flag
Claiming it never makes mistakes.

3. Agent Frameworks

Orchestration. The best answers start with "do you even need an agent?"

  • LangGraph
  • CrewAI
  • AutoGen
  • OpenAI Agents SDK
  • Google ADK
  • Semantic Kernel
When would you use an agent instead of a fixed workflow?
What they are testing
Restraint. Agents cost more and are less predictable.
A strong answer
If the steps are known, use a deterministic workflow with LLM calls inside. Use an agent only when the path depends on intermediate results and cannot be listed upfront. Start simple, add autonomy where an eval shows it helps.
Red flag
Reaching for multi-agent by default.
How would you choose between LangGraph, CrewAI and AutoGen?
What they are testing
Can you map framework style to requirements?
A strong answer
Explicit state, branching, checkpoints and human approval point to a graph model like LangGraph. Role-based task delegation fits crew-style frameworks. Conversational multi-agent fits AutoGen-style setups. Decide on determinism, observability and team familiarity, and prototype both on one real task.
Red flag
Choosing by GitHub stars.
How do you stop an agent looping or burning budget?
What they are testing
Production safety.
A strong answer
Max steps, per-run token and cost budgets, timeouts, tool allow-lists, loop detection on repeated calls, and human approval for destructive actions. Log every step so you can replay a failure.
Red flag
Relying on the prompt saying "stop when done".

4. MCP Layer

The Model Context Protocol connects models to tools and data in a standard way.

  • MCP
  • FastMCP
  • GitHub MCP
  • Filesystem MCP
  • PostgreSQL MCP
  • Slack MCP
What is MCP and why does it matter?
What they are testing
Do you understand the problem, not just the acronym?
A strong answer
An open protocol where a client (the AI app) talks to servers that expose tools, resources and prompts. It replaces one-off integrations per app and per model with one standard interface, so a tool is written once and reused.
Red flag
"It is just function calling." It standardises discovery and transport as well.
What are the security risks of connecting MCP servers?
What they are testing
Threat modelling.
A strong answer
Untrusted servers, over-broad permissions, and prompt injection through tool output. Mitigate with least privilege, read-only by default, authentication, confirmation before writes or deletes, auditing every call, and treating tool results as data, never as instructions.
Red flag
Installing community servers with full database or filesystem access.
MCP server, plain REST API, or in-app function calling?
What they are testing
Trade-off judgement.
A strong answer
Function calling inside one app is simplest. MCP pays off when several clients or models need the same tools. A REST API stays right for non-AI consumers. Often the MCP server is a thin wrapper over an existing API.
Red flag
Rewriting working APIs as MCP for its own sake.

5. Context & Knowledge

Retrieval and vector stores. Expect "why is your RAG bad?" more than "which database?"

  • Pinecone
  • Chroma
  • Weaviate
  • Milvus
  • FAISS
  • OpenSearch
How would you pick a vector store?
What they are testing
Fit to scale, ops capacity and existing stack.
A strong answer
Prototype with something light (Chroma or FAISS). Managed (Pinecone) when you want no ops. Weaviate or OpenSearch when you need keyword plus vector hybrid search or already run the stack. Milvus when you self-host at large scale. Note FAISS is a library: no persistence, auth or filtering out of the box.
Red flag
Naming a favourite with no workload numbers or ops reasoning.
Your RAG answers are poor. What do you check?
What they are testing
Systematic debugging with measurements.
A strong answer
Measure retrieval separately from generation. Build a golden question set and track hit@k and MRR. Then work through chunking, metadata filters, hybrid search, a reranker, and query rewriting, changing one thing at a time. Only then tune the prompt.
Red flag
Immediately swapping the model or the vector database.
When is a graph better than vectors?
What they are testing
Do you add complexity only when proven?
A strong answer
When questions need multi-hop relationships across entities, which plain similarity search handles badly. Prove it on a golden set of multi-hop questions first; otherwise vectors plus a reranker are simpler and cheaper.
Red flag
Choosing GraphRAG because it sounds advanced.

6. Runtime & APIs

Where the code runs. LLM apps stress timeouts, streaming and scaling.

  • FastAPI
  • Docker
  • Kubernetes
  • AWS Lambda
  • Cloud Run
Serverless functions or containers for an LLM backend?
What they are testing
Operational trade-offs specific to LLM workloads.
A strong answer
Functions are cheap and scale to zero but have request time limits and cold starts, which hurts long generations and large model loads. Containers (Cloud Run or Kubernetes) allow streaming, longer jobs and warm models. Use queues and workers for long-running agent tasks.
Red flag
Ignoring timeouts and streaming.
How do you run long agent tasks reliably?
What they are testing
Distributed-systems basics.
A strong answer
Put the task on a queue, run it in a worker, make steps idempotent, persist state between steps, retry with backoff, and report status to the user by polling or events.
Red flag
Holding one HTTP request open for minutes.

7. Evaluation & Testing

The most underrated interview topic. Showing an eval habit sets you apart.

  • LangSmith
  • TruLens
  • DeepEval
  • Promptfoo
How do you test a non-deterministic system?
What they are testing
Do you have a real evaluation practice?
A strong answer
A golden set of representative cases, including ones that must be refused. Deterministic checks where possible (format, citations, tool calls), LLM-as-judge for fuzzy quality after calibrating it against human labels, thresholds in CI, and periodic human spot checks.
Red flag
"I try a few prompts and see if it looks right."
What do you measure for a RAG system?
What they are testing
Separating retrieval quality from answer quality.
A strong answer
Retrieval: hit@k and MRR. Answers: groundedness (is each claim supported by the sources), relevance, and correct refusal when the corpus has no answer. Track cost and latency alongside.
Red flag
One vague "accuracy" number.
How would you gate a release on quality?
What they are testing
Regression discipline.
A strong answer
Run the eval suite on every prompt, model or retrieval change. Fail the build below a recorded floor, and never lower the floor without writing down why.
Red flag
Shipping prompt changes with no regression check.

8. Observability

You cannot fix what you cannot see. Traces, cost and quality in production.

  • Langfuse
  • Helicone
  • OpenTelemetry
  • Grafana
What do you log for an LLM application?
What they are testing
Operational maturity, and privacy awareness.
A strong answer
One trace per request: model, latency, token counts and cost, tool calls, retrieved chunks, errors and user feedback. Redact or avoid storing sensitive prompt text, and keep keys out of logs entirely.
Red flag
Logging full prompts with personal data and no retention rule.
How do you detect quality regressions in production?
What they are testing
Closing the loop after launch.
A strong answer
Sample production traces into the eval pipeline, track feedback signals (thumbs, corrections), alert on cost, latency and error spikes, and compare quality by model or prompt version.
Red flag
Waiting for users to complain.
Why use OpenTelemetry instead of a single vendor SDK?
What they are testing
Lock-in awareness.
A strong answer
A standard lets you change or add backends (Langfuse, Grafana, others) without re-instrumenting the code, and it correlates AI spans with the rest of your services.
Red flag
Not knowing what it is.

9. Deployment

Shipping and rolling back AI features safely, at a cost you can predict.

  • Vercel
  • AWS
  • Azure
  • Google Cloud
How do you deploy and roll back an AI feature?
What they are testing
Release engineering for non-deterministic code.
A strong answer
Version prompts and models as config, gate releases on the eval suite, roll out behind a feature flag or canary, watch cost and quality metrics, and keep a one-step rollback to the previous prompt or model.
Red flag
Editing a production prompt in place.
How do you keep AI cost predictable?
What they are testing
Do you treat cost as a requirement?
A strong answer
Per-user and per-tenant limits, output caps, caching, routing simple steps to cheaper models, and budget alerts. Measure cost per completed task, not just token price.
Red flag
No limits, so one abusive user sets your bill.
Which cloud, and why?
What they are testing
Pragmatism over loyalty.
A strong answer
Start from constraints: where your data and team already are, compliance needs, and which managed model services you need. Name the lock-in you accept and the abstraction that limits it.
Red flag
Naming a cloud with no constraint behind the choice.