Gradion Lead Senior Fullstack Interview Preparation
4 October 2026
How to use this handbook
This course prepares Tony for the hands-on Lead/Senior Fullstack individual contributor role in the supplied Gradion brief. The target is to explain mechanisms below frameworks, demonstrate ownership of products with real users, defend engineering trade-offs, and show where AI earns its place in a workflow. Use the short answers aloud, then defend them against the probes. Code is illustrative unless an exercise explicitly says it was run. The study schedule assumes 14 days at 90 minutes a day because an interview date and time budget were not provided.
First move: prepare one evidence-rich CoverGo product story, one InvestaX reliability story, and one AI evaluation story. Then do the technical diagnostic below. A stack inventory is not a proof of senior judgment.
Evidence and scope
| Evidence | What it supports | Limits |
|---|---|---|
| Supplied four-page Gradion requirements PDF, accessed 4 October 2026 | Hands-on IC, architecture, reliability, AI APIs/agents/RAG/evaluation, client communication, mentoring, product end users, deep CS and framework internals | Interview rounds and specific client stack are not given. Its experience threshold is 5+ years. |
| Supplied three-page CV, accessed 4 October 2026 | 6+ years claimed; CoverGo distribution, Temporal and underwriting AI; InvestaX trading/migration; Select IoT and legacy modernization; Hitachi GIS and fleet work | CV claims are candidate statements. Metrics, attribution, versions and production status need supporting evidence. |
| Gradion’s current Lever posting, checked 4 October 2026 | Closely matching role, explicit AI hands-on requirement, depth in one language, 7+ years | The live posting differs from the supplied PDF; recruiter should clarify which threshold and role variant applies. |
| Gradion careers, full-stack services, and AI practice, checked 4 October 2026 | Client-embedded engineering, heterogeneous stacks, end-to-end product work, measurable AI outcomes | Company-wide descriptions do not prove a particular team’s architecture or interview format. |
The 5+ versus 7+ experience discrepancy is material. The CV says 6+ and lists June 2019 to present, which is over seven calendar years by October 2026, but the stated figure and dates differ. Reconcile the actual full-time dates and the summary wording before an interview; do not improvise. Likewise, distinguish an underwriting agent built and evaluated from one fully integrated into production: the CV says integration into the core platform is in progress.
Company and role brief
Gradion is a technology consulting and embedded engineering business serving clients in commerce, travel, fintech and other domains. Its public descriptions emphasize product delivery and recovery, full-stack work across several languages, and AI tied to measurable cost or revenue outcomes. The Shopware case study reports a 21-person embedded team and AI Co-pilot features; the HomeToGo case study reports a long-lived engineering partnership and many external integrations. These are public case-study claims, not a map of the team Tony would join. A strong interview answer moves from client problem to constraints, decision, evidence, rollout and measured result. Company sources.
Likely rounds, inferred rather than confirmed: recruiter and role fit; technical coding/debugging; system design; production/AI deep dive; stakeholder and peer discussion. Ask the recruiter for actual format, time, primary language, and whether a live codebase is used. The supplied JD does not require PHP, C/C++, AWS, Go, React or a specific LLM vendor. Use Go and TypeScript/React as Tony’s evidence-bearing languages while being ready to reason in an unfamiliar stack.
Requirement to preparation map
| Priority | Requirement | CV evidence to prepare | Gap or probe | Sections |
|---|---|---|---|---|
| Must | CS depth beneath framework; language choice | Go workers/services, Java/OSGi legacy, Node services | Explain memory, event loop, I/O, profiling, and a framework-free HTTP request | 1, 2, Q01-Q18, E01-E04 |
| Must | Architecture, scale, reliability, debt | InvestaX migration and pool tuning; CoverGo tenancy; Select modernization | Baselines, alternatives, failure modes and measured outcomes | 3, 5, 7, D01-D03 |
| Must | Hands-on product and end-user feedback | CoverGo quote/policy/channel; Select Aura and IoT; InvestaX trading | Identify actual users, feedback loop, release aftercare and personal contributions | 8, story bank |
| Must | AI APIs, agents/RAG/evaluation | Underwriting agent, MCP, Qdrant/pgvector, eval harness | Show dataset, scoring, false decisions, permissions, cost and production boundary | 6, D02, E11 |
| Must | Client-facing technical direction and influence | Platform fundamentals, reusable SDKs, cross-team tools | One disagreement changed by evidence, and one decision to say no | 8, story bank |
| Supporting | Frontend and complete delivery | React/Vue dashboards and end-to-end modules | Accessibility, data loading, error/empty states and field feedback | 4, E08 |
| Supporting | Automation/DevOps/LiveOps | CI/CD, K8s, monitoring, release cadence | A real incident timeline and rollback | 5, D04 |
| Optional until team stack known | Deep C/C++, PHP and vendor-specific services | No detailed proficiency established by CV | Demonstrate fundamentals and fast learning; avoid claiming specialist depth | follow-up study |
Quick diagnostic
Without notes, take 25 minutes to: (1) explain how a browser request reaches a Go or Node handler and a PostgreSQL row; (2) show how two concurrent requests can both issue the same order; (3) identify a slow endpoint using traces, pool metrics and an execution plan; (4) describe an AI golden set with false-approval costs; (5) narrate one user feedback loop and one technical disagreement. Score each from 0 absent, 1 fragmentary, 2 workable with gaps, 3 correct with trade-offs, 4 robust under follow-up. Any score under 3 becomes the first study block.
1 Computing foundations
From click to result
A browser resolves a name, establishes a transport connection and TLS when applicable, sends an HTTP request, and may reuse the connection. A proxy routes it; the process parses headers and body, authenticates, authorizes, validates, calls domain code, obtains a database connection, runs a transaction or query, and serializes the response. Every queue and remote hop adds latency; network waits, lock waits and connection-pool waits do not necessarily consume CPU. Follow one trace across boundaries before blaming the framework. At the bottom, the runtime schedules work onto OS threads; syscalls cross into the kernel; cache locality, allocation and synchronization shape cost. HTTP overview; Node event loop; Go diagnostics.
Mental model: throughput is completed requests per second, latency is elapsed time for one request, and concurrency is work in flight. Little’s Law, L = lambda W, lets a team estimate average in-flight work from stable arrival rate and mean response time. At 200 requests/s and 0.5 s mean, about 100 requests are in flight. Averages hide tails: compare p50, p95 and p99 under load, and specify where they are measured. A faster CPU will not repair an exhausted connection pool or a slow external dependency.
Data structures: array/slice indexing is O(1) and tends to favor locality; linked traversal is O(n) and pointer-heavy; a hash table gives expected O(1) lookup with hashing and resizing costs; a balanced tree supports ordered O(log n) operations. Choose on access pattern, size, contention and measured memory, not notation alone. A DB B-tree has different storage and I/O behavior from an in-process tree. Practice bounded queues and two-pointer scans before esoteric puzzles.
Debug from evidence
Draw a latency budget across DNS/TLS, gateway, service CPU, pool wait, SQL, downstream calls and response. Use request ID/trace ID, dependency timings, error ratio, queue depth and saturation. Distinguish wall-clock profiles from CPU profiles: a CPU profile can be quiet when goroutines or callbacks are waiting on a pool. Reproduce with a controlled workload, state a hypothesis, change one variable, and compare a before/after distribution. The post-fix test must include the failure boundary that originally broke. OpenTelemetry signals; Go profiling.
2 Language and runtime mechanics
Go depth
Go values have concrete representation and ownership conventions. Slices are descriptors over backing arrays: appending may allocate, while subslices can retain a large array. Maps are reference-like but not safe for unsynchronized concurrent writes. Goroutines are cheap relative to OS threads, never free; bound them by queue capacity, semaphores and cancellation. A channel coordinates data flow; a mutex protects shared state; neither should be chosen by slogan. Data-race freedom gives stronger reasoning about visibility under the Go memory model. Use go test -race on representative concurrent tests, then CPU/heap/block/mutex profiles as the symptom dictates. Pass context.Context down request lifetimes and honor deadlines; do not store request contexts for detached work. Race detector; Diagnostics.
TypeScript and Node depth
TypeScript checks source types; JSON and external inputs still require runtime validation. Node runs JavaScript callbacks on the event loop while I/O and some work are delegated; a long synchronous parse, regex or crypto operation can starve other requests. async creates a Promise-based flow, not a new CPU thread. Set bounded concurrency, cancellation and explicit error propagation for remote calls. A rejected Promise without correct handling can strand a request or create a process-level error depending on runtime and configuration. Know the deployed Node version; do not rely on one microtask/timer example as a universal event-loop explanation. Node guidance.
Choosing the tool
For a small product team, TypeScript end to end can reduce context switching; Go can suit a compact service with high concurrency and simple deployment; Java may suit an established transactional ecosystem; Python can accelerate data-adjacent work. Decide from workload, team expertise, ecosystem, operability and migration cost. Prove a bottleneck with a representative load test before a language rewrite. Gradion’s full-stack page names several languages and explicitly ties choice to the problem. That is a company position, not evidence that every client uses each language.
3 Data and distributed systems
Transactions, invariants and indexes
Start with the business invariant, such as one accepted quote producing at most one active policy per insurer reference. Enforce the narrowest durable invariant with a unique constraint and a transaction; use a request idempotency key for replayed commands. PostgreSQL Read Committed uses a fresh snapshot per statement, so a check-then-insert can race. A unique constraint may suffice; more complex cross-row invariants may need a lock, careful update predicate, or Serializable plus whole-transaction retry. Retry serialization failures only when side effects are outside the transaction or safely idempotent. Transaction isolation.
Index columns matching filtering, joins and ordering, then inspect EXPLAIN (ANALYZE, BUFFERS) on realistic data. A composite index on (tenant_id, status, created_at DESC) can serve tenant/status filtering and recency, subject to query shape and selectivity. Indexes consume storage and slow writes. A sequential scan may be correct for a small table or broad filter. Pool size must reflect actual DB capacity across replicas and workers: 20 connections times 30 pods is 600 potential sessions, not 20. Pool wait, active/idle connections, DB CPU and lock waits tell different stories. Using EXPLAIN; Examining index usage.
Boundaries and delivery
Keep a modular monolith if a small team benefits from one transaction and deployment. Split when domain ownership, independent scaling or deployment pressure justifies the new failure modes. A synchronous API gives an immediate contract but couples latency and availability; a queue decouples work but creates backlog, ordering and duplicate-delivery concerns. The outbox stores business state and an event in the same database transaction; a publisher then delivers at least once, and consumers deduplicate by event or operation ID. An outbox is not atomic delivery to a broker. Cache only with a defined freshness contract and invalidation path. For tenant data, put tenant scope into authorization and query predicates; test cross-tenant access and avoid trusting a client-supplied tenant ID. AWS retry guidance.
4 Frontend and user experience
The UI is part of correctness. Separate server data from local interaction state, keep a clear loading/empty/error/success state, and guard stale responses when selection changes quickly. React effects synchronize with external systems; deriving render data through an effect often creates an avoidable update loop. Measure before adding memo or useMemo; they are optimizations, not semantic guarantees. For large tables, paginate or virtualize with accessibility in mind. Track Largest Contentful Paint, Interaction to Next Paint and Cumulative Layout Shift using field data, then investigate the actual bottleneck. A keyboard-only journey through quote and policy screens catches failures that a screenshot misses. React useEffect; React memo; Web Vitals.
5 Production engineering
Reliability and release discipline
Define an SLI from a user journey, not simply container health: e.g. eligible quote-to-policy completion without duplicate issuance within an agreed time. Set an SLO with a window and exclusions; the error budget makes release-versus-reliability choices explicit. Instrument traces, metrics and structured logs with stable identifiers but avoid recording sensitive policy data. During an incident, mitigate first, communicate impact and uncertainty, preserve evidence, then find the causal chain and write a durable action. A deployment should have small increments, feature flags or canaries, schema compatibility, and a rollback route. Google SRE error budgets; OpenTelemetry signals.
Review, tests and debt
Test pure domain invariants quickly, then exercise transaction boundaries and external contracts, plus a few end-to-end user journeys. A mock of a repository returning the expected row cannot prove that concurrent SQL writes are safe. Review risky assumptions, data flow, permissions, retries and migrations before style. For technical debt, name the user/business cost, collect a baseline, choose a small measurable slice and set a stop rule. For an unfamiliar codebase, trace one vertical user journey, run tests, inspect production signals and make a reversible change before proposing a rewrite.
6 AI systems that earn their cost
The decision ladder
Begin with deterministic rules or search when the problem is well specified. Use an LLM when unstructured input or language variation creates real value. RAG retrieves authorized, versioned source chunks, then generates an answer with citations and abstention when evidence is missing. An agent adds tool selection and iteration; use a bounded workflow for known steps. Keep system instructions separate from retrieved text; treat documents and tool output as untrusted. Check tool arguments and resource authorization in code, not in a prompt. Gate consequential writes with a human or a narrow deterministic policy. Anthropic tool use; prompt injection guidance.
Evaluate from the failure cost
For underwriting, define cases by product, rule version, language, document quality, ambiguous exclusions and adversarial content. Freeze a golden set with expert adjudication and separate a held-out set. Measure retrieval recall and citation support separately from final decision correctness; report false approvals, false denials, abstention, latency, cost and escalation rate by slice. Use code grading for exact constraints, expert review for consequential judgments, and calibrated model grading only where reliable. Re-run after prompt, model, retrieval or policy changes; monitor drift and feedback in production. State the business cost of a false approval versus a manual review. Anthropic evaluation guidance; Anthropic agent evals.
Workflow and protocol boundaries
Temporal Workflow code must replay deterministically; LLM calls belong in retryable Activities, with idempotent external effects. LangGraph checkpoints agent state and supports human interrupts; choose it for agent orchestration when that state model helps, not as a replacement for all business transactions. MCP exposes tools/resources through a protocol; it does not make a tool authorized. Scope credentials and enforce per-tool, per-tenant policy server-side. The MCP specification changed in July 2026; confirm the deployed protocol and SDK version before naming transport details. Temporal workflow definition; LangGraph JavaScript; MCP 2026 release.
7 Architecture and technical direction
Start system design with actors, outcomes, invariants, volume, peak shape, data classification, latency and recovery targets. Draw the smallest design that meets them. Estimate units explicitly: 100k daily commands divided by 86,400 seconds is about 1.16/s average, but a 20x peak and downstream fan-out are the meaningful sizing inputs. Decide where a transaction ends, who owns data, which calls can fail, and how a client sees partial progress. Discuss capacity and operational cost, then migration and rollback. Write a short decision record: context, options, decision, consequences, trigger to revisit. Influence comes from a prototype, benchmark, review and shared ownership rather than an architecture decree. Gradion architecture review.
8 Product, client and influence
When a client asks for an AI feature, learn the user journey, current error and time cost, available data, acceptable error, consent/retention needs and who will own exceptions. Offer a cheap baseline, a bounded pilot and a kill criterion. When a user says the dashboard is slow, translate it into the exact journey, device, dataset and perceived delay, then measure real-user impact. Close the loop after release: contact users, compare before/after measures and fix new failure modes. A senior IC can say no with evidence and an alternative, record the decision, and remain accountable for delivery. Gradion explicitly emphasizes product/business acumen, critical thinking and communication in its careers framework.
Answered interview questions
All questions below are curated practice for this role, not claims about Gradion’s actual interview bank. M means must master; S means supporting. Say the answer in 30-90 seconds, then use the probe to extend it. A strong response states an invariant or measurement, an alternative, and a failure mode. The linked primary source supports the technical mechanism.
Computing and debugging Q01 to Q08
Q01 M Core What happens from URL to rendered page?
Tests: layered reasoning. Answer: DNS resolves the host; transport and TLS establish a protected connection; HTTP requests cross proxies to an application, which authenticates, authorizes, calls domain/data services, and returns bytes; the browser parses, loads resources, builds and paints the page. Connections may be reused and caches may short-circuit steps. Probe: A slow page with a fast API could still suffer render-blocking assets or client computation. Name network and rendering measurements separately. Trap: Saying that every request creates a new TCP connection. HTTP.
Q02 M Core Why might CPU be low while p95 latency rises?
Tests: bottleneck diagnosis. Answer: Work may wait for a DB connection, lock, queue, remote service or rate limit. Compare trace span durations, pool wait, queue length, DB waits and saturation; CPU profiles only sample active CPU use. Probe: Load 200/s at 500 ms average implies roughly 100 in-flight operations if stable; as service time rises, concurrency and queueing can spiral. Trap: Adding replicas without checking shared DB capacity. Tracing.
Q03 M Core Explain cache locality versus big-O.
Tests: below-framework performance. Answer: Contiguous arrays often exploit spatial locality and fewer allocations, while pointer-chasing structures suffer cache misses even with similar asymptotic complexity. A hash lookup is expected O(1), yet hashing, collision and memory footprint matter. Benchmark representative input and look for allocation or cache miss evidence before replacing data structures. Probe: For ordered range scans, a tree or sorted slice may win depending on update frequency. Trap: Treating O(1) as always faster than O(log n). Go diagnostics.
Q04 M Core Stack, heap, and garbage collector?
Tests: memory model. Answer: Call frames and local values may live on a goroutine stack, while escaping objects require longer-lived allocation; Go can grow stacks and its compiler performs escape analysis. The GC traces reachable heap objects and consumes CPU and memory headroom. Profile retained objects and allocation rate, not just heap size. Probe: A small slice of a huge backing array can keep the array live. Trap: Asserting every local variable lives on the stack or manually freeing managed objects. Go diagnostics.
Q05 M Core Process versus thread versus goroutine?
Tests: scheduling. Answer: A process has its address space and resources; OS threads are scheduled execution contexts sharing a process’s memory; Go goroutines are runtime-managed tasks multiplexed over threads. Goroutines still allocate stack and scheduling state, and blocked work can saturate downstream resources. Probe: What limits 100k goroutines all waiting on a DB? Memory and pool/DB capacity, so bound in-flight work. Trap: Confusing goroutine count with throughput. Go runtime.
Q06 M Core What is the difference between a timeout and cancellation?
Tests: failure propagation. Answer: A timeout sets a deadline for a particular operation; cancellation signals that work is no longer needed. Propagate a remaining deadline through remote calls and stop expensive work when the caller disappears, but make cleanup safe. A timeout does not prove the remote side failed: it may have committed and the response was lost. Probe: After a payment timeout, query by idempotency key before replaying. Trap: Blindly retrying every timeout. AWS timeouts.
Q07 S Core Compare latency, throughput, concurrency and saturation.
Tests: capacity vocabulary. Answer: Latency is duration per operation, throughput completed rate, concurrency work in flight, saturation the degree a finite resource is occupied or queued. At increasing offered load, throughput eventually flattens and latency rises sharply. Report percentiles and workload shape, not only mean. Probe: A 99.9% availability objective over 30 days permits about 43.2 minutes of unavailable time if uniformly measured; user-facing request SLIs may calculate a different error budget. Trap: Equating utilization with user success. SLOs.
Q08 S Core How do you approach an unfamiliar codebase in a day?
Tests: practical learning. Answer: Reproduce one user path locally, map entry point to state/data and deployment, read tests and recent incidents, then make a tiny reversible change with a safety test. Inspect logs and traces to understand hot paths and ask the owner about invariants. Document findings and unknowns. Probe: A broad static diagram is weaker than tracing a concrete quote-to-policy request. Trap: Proposing a rewrite before discovering what users and jobs the system serves. Gradion role.
Go and TypeScript Q09 to Q18
Q09 M Go What happens when you append to a slice?
Tests: aliasing. Answer: A slice holds pointer, length and capacity. Append reuses the backing array if capacity permits; otherwise it allocates and copies. Two slices may therefore share mutations until growth separates them. Copy when ownership must be isolated, and avoid retaining a small view of a huge buffer. Probe: Predict whether a := make([]int,1,2); b := a[:2]; a=append(a,9) changes b[1]: yes, capacity allowed reuse. Trap: Assuming append always mutates the original variable’s storage. Go specification.
Q10 M Go When use a mutex rather than a channel?
Tests: concurrency design. Answer: A mutex is straightforward for a short critical section protecting shared state; channels suit ownership transfer or ordered communication. Minimize lock scope and prove there is no unlocked access. Neither form gives business-level exactly-once across processes. Probe: For counters on a hot path, compare atomic operations, sharding and mutex contention with a benchmark. Trap: “Always use channels” as a rule. Go memory model.
Q11 M Go How do you find a data race?
Tests: evidence. Answer: Run concurrent tests with go test -race, reproduce representative interleavings, inspect shared reads/writes, and fix ownership or synchronization. The detector only reports executed paths; a clean run is not a proof. Retest under load and add an invariant test. Probe: A concurrent map write may fail; a read/write without synchronization is also unsafe even if the program seems to work. Trap: Treating a channel somewhere in the function as blanket synchronization. Race detector.
Q12 M Go How should context flow through a service?
Tests: lifecycle. Answer: Accept a context at the request boundary and pass it to DB and remote calls; derive tighter deadlines where justified, defer cancel, and stop work on cancellation. Background jobs need their own explicit lifetime rather than retaining a request context. Probe: A downstream SDK that ignores cancellation can still occupy resources; set its client timeout and bound concurrency. Trap: Calling context.Background() in an inner helper and silently severing cancellation. Go context.
Q13 S Go Which profile for rising memory and slow requests?
Tests: profiling choice. Answer: Heap profile for retained allocations, allocation profile for churn, CPU profile for active cycles, block/mutex profiles for waiting or contention, and trace for scheduling and latency. Correlate with GC, pool wait and request distribution. Probe: High allocations can hurt via GC even when the retained heap is flat. Trap: Inferring a memory leak from RSS alone. Go diagnostics.
Q14 M Node What does async buy you and what does it not?
Tests: event loop. Answer: async/await expresses asynchronous completion and error propagation through Promises. It does not move CPU-heavy JavaScript to another thread. Nonblocking I/O lets one loop manage many outstanding operations, provided callbacks stay short. Probe: A large synchronous JSON transform blocks every request on that event loop; use batching, workers or another service after measurement. Trap: Saying Node executes all JavaScript in parallel because it uses libuv. Node guidance.
Q15 M Node How do you bound 10k outgoing calls?
Tests: backpressure. Answer: Use a bounded worker pool or semaphore, define per-call timeout and overall deadline, and propagate cancellation. Queue only within a limit; reject or defer excess load. Record in-flight count, queue age and failure ratio. Probe: Promise.all eagerly starts all mapped requests, so it does not bound concurrency. Trap: Set a huge pool and push saturation to the dependency. AWS retry guidance.
Q16 M TypeScript Why validate DTOs at runtime?
Tests: trust boundary. Answer: TypeScript types disappear at runtime; JSON, headers and persisted data can violate them. Parse and validate at input boundaries, reject unknown or invalid values deliberately, then use typed domain objects. Validate output contracts in integration tests too. Probe: An enum widened by a partner API should not silently become a default financial decision. Trap: as Policy does not validate a payload. TypeScript handbook.
Q17 S Runtime Compare Go and Node for an I/O-heavy service.
Tests: judgment. Answer: Both can handle concurrent I/O; Go uses goroutines and a compiled runtime, Node uses an event-loop model with a broad JS ecosystem. Choose based on team, libraries, deployment, observability, data contracts and measured performance. CPU-heavy work and process isolation may change the answer. Probe: Describe an experiment with representative request mix, p95, memory and operational effort. Trap: Claiming a universal throughput winner from language identity alone. Gradion full-stack.
Q18 S Runtime What does a framework-free handler still need?
Tests: fundamentals. Answer: Parse and limit input, route method/path, validate and authorize, apply domain logic, manage transactions, map errors to HTTP, set timeouts, record telemetry and shut down cleanly. A framework packages conventions and middleware but does not remove those obligations. Probe: Implement a tiny net/http endpoint that enforces body size and context deadline. Trap: Mistaking fewer layers for no security or lifecycle work. Go net/http.
Data and distributed systems Q19 to Q30
Q19 M Data How do you prevent a double-issued policy?
Tests: invariant and concurrency. Answer: Define the uniqueness key from the business rule, enforce it with a database constraint, wrap state transition in a transaction and store an idempotency key/result for command retries. Return the same result on a duplicate key when the payload matches; reject a conflicting payload. Probe: A crash after commit but before response is the central replay case. Trap: A read-then-insert check alone races across replicas. PostgreSQL isolation.
Q20 M Data Read Committed versus Serializable?
Tests: isolation. Answer: In PostgreSQL, Read Committed snapshots each statement; Serializable admits only results consistent with a serial order but can abort a transaction requiring full retry. Choose the least costly mechanism that protects the invariant: unique constraints, conditional updates or row locks may be simpler. Probe: Do not repeat an irreversible external call inside a retried transaction. Trap: Calling Read Committed ‘no consistency’ or Serializable ‘zero retries.’ PostgreSQL isolation.
Q21 M Data Why might PostgreSQL ignore an index?
Tests: query plans. Answer: The planner estimates a sequential scan cheaper for small or low-selectivity data, or the index may not match predicates/order, statistics may be stale, or an expression/cast changes usability. Inspect EXPLAIN (ANALYZE, BUFFERS) and row-estimate errors with production-like data. Probe: Consider write cost before adding a composite index to every filter. Trap: Forcing index use without reading the plan. EXPLAIN.
Q22 M Data Why does scaling pods exhaust PostgreSQL?
Tests: connection budgets. Answer: Each process may own a pool; replicas times workers times max pool is the potential connection count. Active transactions, locks and query time make each connection scarce. Measure pool wait and DB capacity, lower per-pod limits, add admission control or a suitable proxy, and fix long transactions. Probe: Under 30 pods and a pool max of 20, potential 600 sessions; confirm shared quotas and failover behavior. Trap: Increasing max connections as the only remedy. PostgreSQL runtime settings.
Q23 M Distributed What does an outbox solve?
Tests: dual write. Answer: Commit the domain state and an event record atomically in one DB transaction. A publisher later sends the event; it can send twice, so consumers deduplicate and handlers are idempotent. The outbox closes the DB/broker dual-write gap but adds lag, cleanup and monitoring. Probe: If a publisher crashes after broker acknowledgment but before marking sent, duplicate delivery is expected. Trap: Claiming exactly-once end-to-end. AWS transactional outbox guidance.
Q24 M Distributed REST, gRPC or event?
Tests: contract choice. Answer: REST is convenient for external resource APIs; gRPC offers typed contracts and efficient service calls; events decouple temporal availability when callers can tolerate eventual completion. Choose per consumer, failure semantics and ownership, not speed alone. Document errors, versioning, deadlines and idempotency. Probe: A quote decision needing immediate user feedback may use a synchronous command while document processing emits a completion event. Trap: Replacing a database transaction with a chain of remote calls. HTTP methods.
Q25 M Distributed Cache invalidation strategy?
Tests: freshness. Answer: State the stale-data tolerance per object. Use cache-aside with TTL for read-heavy noncritical data, versioned keys or event invalidation where necessary, and a source-of-truth fallback. Negative caching and stampede protection matter at scale. Keep authorization and tenant boundaries in keys. Probe: A revoked policy permission may require immediate authoritative check rather than a five-minute stale cache. Trap: Caching a tenant-independent object that contains tenant-scoped data. AWS caching guidance.
Q26 M Distributed How do you handle retries safely?
Tests: failure mode. Answer: Retry only transient failures within a deadline, with capped exponential backoff and jitter; make mutation operations idempotent and define a key or sequence. Budget retries at one layer to avoid multiplicative storms. Distinguish a timeout from a failed commit. Probe: A 3-layer stack each making three attempts can produce 27 calls if uncontrolled. Trap: Retrying non-idempotent writes because the HTTP status was 5xx. AWS retries.
Q27 M Distributed Why start with a modular monolith?
Tests: architecture restraint. Answer: It preserves one deployment and transactions, while domain modules can still have clear interfaces and tests. Move to services when team ownership, independent scaling or release cadence justifies network failure and data duplication costs. Probe: Extract by business capability with an anti-corruption boundary and gradual traffic shift. Trap: Treating microservices as a maturity badge. Gradion full-stack.
Q28 M Data How do you enforce tenant isolation?
Tests: security. Answer: Resolve tenant from trusted identity context, authorize the action and resource, and constrain every data access. Separate public/admin paths, test negative cross-tenant cases and audit background jobs. Row-level security can add defense in depth if context and policies are managed correctly. Probe: A cache key, search index or event consumer may leak even when SQL filters are correct. Trap: Trusting tenant_id from the request body. PostgreSQL row security.
Q29 S Data How do you migrate a large table safely?
Tests: deployment compatibility. Answer: Add an optional field/index first, deploy dual-readable code, backfill in batches while measuring locks/replication lag, switch writes/reads, then remove old shape after verification. Avoid a giant transaction and ensure old code can coexist during rollback. Probe: Concurrent index creation and validation have version-specific constraints; verify the deployed PostgreSQL version. Trap: A single deploy that requires all pods to upgrade atomically. PostgreSQL indexes.
Q30 S Distributed What is backpressure in an event consumer?
Tests: queue health. Answer: Bound in-flight messages and downstream calls; measure lag, age and retry/dead-letter volume. If processing capacity falls below arrivals, scale within downstream limits, shed or defer work, and prioritize important traffic. Preserve partition ordering where the business needs it. Probe: A fast broker cannot make a slow database faster. Trap: Increasing consumer concurrency without considering locks and pool connections. AWS reliability.
Frontend Q31 to Q38
Q31 M Frontend Why does a React component render again?
Tests: rendering mental model. Answer: State or parent/context updates can trigger rendering; React reconciles output and may change only part of the DOM. A render is not identical to a paint. Profile a real slow interaction before memoizing and reduce unnecessary state coupling. Probe: memo can skip some prop-stable renders but is not a correctness guarantee. Trap: Treating every render as a full DOM rebuild. React memo.
Q32 M Frontend When should you use an Effect?
Tests: state design. Answer: Use it to synchronize with an external system such as a subscription, network or DOM API; include cleanup for subscriptions and cancellation or stale-response guarding for fetches. Compute derived display values during render when possible. Probe: An effect that sets state from props may create extra renders and races. Trap: Putting all business logic into effects. React useEffect.
Q33 M Frontend How do you stop a stale search response?
Tests: race reasoning. Answer: Key the response to the current query/request or abort the previous fetch; ignore results from outdated requests. Server APIs should still be safe to receive both. Show loading, empty and error states and preserve user intent on retry. Probe: A canceled browser fetch does not guarantee the server canceled an already-started DB transaction. Trap: Assuming response order equals request order. React useEffect.
Q34 M Frontend Server state versus UI state?
Tests: data ownership. Answer: Server state is asynchronously loaded, shared and potentially stale; a query cache needs keys, invalidation and refetch policy. Local UI state describes current selection, open panels or draft input. Keep policy decisions on the server and avoid duplicating server records in several stores. Probe: An optimistic accepted-quote update must reconcile a rejected server result. Trap: Using a global store as the single authority for a financial workflow. React state guide.
Q35 S Frontend How would you fix a 50k-row dashboard?
Tests: end-to-end performance. Answer: Push filtering/pagination to an indexed API, return only needed columns, cache carefully, then virtualize rendering where appropriate. Measure API p95, payload size, main-thread work and interaction latency. Consider keyboard/screen-reader behavior across virtualized rows. Probe: A faster React component cannot repair an unbounded SQL scan. Trap: Downloading all rows and adding useMemo as the entire fix. Web Vitals.
Q36 S Frontend Which user-facing performance metrics matter?
Tests: measurement. Answer: LCP reflects loading, INP responsiveness and CLS visual stability. Combine field distributions with task-specific metrics such as quote form completion and error rate. Segment by device/network and actual end users. Probe: A lab Lighthouse score is useful for diagnosis but is not the same as field p75. Trap: Reporting only backend p95 for a slow frontend. Web Vitals.
Q37 S Frontend How do you make a form accessible?
Tests: product quality. Answer: Use semantic labels, keyboard order, visible focus, associated error messages, appropriate input types and clear status announcements. Validate with keyboard and assistive technology, not color alone. Keep server validation authoritative. Probe: On submit, direct focus to the error summary or first invalid field without trapping it. Trap: Treating aria-label as a replacement for coherent markup. WAI forms.
Production Q39 to Q46
Q39 M Production SLI, SLO and error budget?
Tests: reliability contract. Answer: An SLI measures a user-relevant outcome; an SLO is a target over a window; the error budget is allowable bad events or time. Choose a precise numerator/denominator and discuss risk when burn accelerates. For a quote journey, success might require a timely valid decision, not merely HTTP 200. Probe: If a dependency is down but cached quotes remain available, measure actual user outcome. Trap: Confusing uptime of one pod with product reliability. Google SRE.
Q40 M Production How do you respond to a p95 incident?
Tests: incident leadership. Answer: Establish impact and timeline, mitigate via rollback/traffic shaping or dependency isolation, communicate uncertainty, correlate traces with DB/pool and release changes, then preserve evidence. After recovery, write a causal postmortem with ownership and a verification date. Probe: A pool wait spike after scaling may be caused by total connections or slow SQL, so inspect both. Trap: Tuning code during an outage before reducing user impact. SRE workbook.
Q41 M Production Safe schema and app rollout?
Tests: reversibility. Answer: Use expand, migrate, contract: compatible schema first, then code that tolerates old/new data, backfill, feature switch, later cleanup. Monitor correctness and performance at each step. Rollback plans include old binaries coexisting with schema changes. Probe: A dropped column cannot be restored by redeploying code. Trap: A one-step destructive migration tied to a release. PostgreSQL ALTER TABLE.
Q42 M Production What tests prove a critical invariant?
Tests: test boundaries. Answer: Unit test domain cases, integration test unique constraints and transaction behavior with a real DB, contract test external consumers, and run a small end-to-end user journey. Include duplicate requests and failure injection. The cheapest test that catches a risk is preferred. Probe: Make concurrent requests with the same idempotency key and assert one durable policy and consistent replies. Trap: Mocking away the exact concurrency behavior under test. PostgreSQL isolation.
Q43 M Production What do you review in an AI-generated PR?
Tests: AI-augmented judgment. Answer: Trace requirements to changes, inspect auth/data boundaries, dependency/API versions, failure behavior, edge cases, tests and operational impact. Run tests and inspect the diff for unrelated changes. The human owner is accountable for the result. Probe: Ask the assistant to propose adversarial tests, then independently verify them. Trap: Treating plausible code and a green happy-path test as proof. Gradion AI practice.
Q44 S Production What belongs in a technical decision record?
Tests: influence. Answer: Record the problem, constraints, options, evidence, decision, consequences, owner and revisit trigger. A short benchmark or spike may be more useful than a long opinion. Make it legible to client and maintainers. Probe: If the assumption changes, the record makes a reversal intelligible rather than political. Trap: Recording only the chosen technology without alternatives. Gradion role.
Q45 S Production When pay down debt?
Tests: prioritization. Answer: Tie debt to release delay, incidents, defect rate, security or operational cost, then choose a bounded change with a baseline and a measurable exit criterion. Defer cosmetic rewrites when user value is low. Probe: If deployment takes two days because of fragile coupling, isolate one seam and compare lead time after two releases. Trap: Calling every disliked legacy design debt. Gradion modernization.
Q46 M Production How do you make observability useful?
Tests: causality. Answer: Correlate logs, traces and metrics by operation, tenant-safe identifiers and release version; instrument business outcome and dependency timings. Keep cardinality bounded and redact sensitive payloads. Set alerts on user-visible symptoms and burn, then dashboards for diagnosis. Probe: Trace sampling can miss rare failures, so retain error exemplars or targeted sampling where policy allows. Trap: Logging complete documents for debugging. OpenTelemetry.
AI Q47 to Q56
Q47 M AI When is RAG appropriate?
Tests: problem selection. Answer: Use retrieval when answers depend on a changing corpus and must cite source passages. If a rule is deterministic and structured, use a rules engine or SQL. Build a baseline search, chunk by document structure, preserve version and ACL metadata, and measure retrieval separately from generation. Probe: Missing evidence should cause abstention or escalation, not invented policy text. Trap: Treating larger embeddings as guaranteed accuracy. Anthropic evals.
Q48 M AI How do you evaluate underwriting decisions?
Tests: risk-sensitive measurement. Answer: Define expert-labeled cases by product and rule version, including ambiguous and adversarial examples. Score false approvals/denials and abstention separately; measure citations, tool actions, latency, cost and human escalation. Hold out cases from prompt iteration and compare to a non-LLM baseline. Probe: If a false approval is more costly than review, tune for safe escalation rather than raw accuracy. Trap: Reporting one aggregate accuracy number. Anthropic evals.
Q49 M AI Agent or fixed workflow?
Tests: autonomy judgment. Answer: A fixed workflow is easier to test when steps are known; an agent is useful when tool choice or exploration varies. Bound tools, budget, steps and permissions, and record traces. Keep final side effects behind deterministic checks or human review. Probe: An underwriting eligibility decision may use a fixed rule path with an LLM extracting evidence; no general agent is required. Trap: Equating autonomy with business value. Anthropic tools.
Q50 M AI What is a prompt-injection boundary?
Tests: adversarial design. Answer: A retrieved PDF or tool result is data from a lower-trust source. It may contain instructions to change behavior; the application must preserve instruction precedence and restrict tool actions with code-level authorization. Test hostile documents, URLs and responses in evals. Probe: If a source asks the agent to email customer data, the tool should be unavailable or deny the operation regardless of prompt wording. Trap: “Tell the model to ignore injections” as the sole control. Anthropic guardrails.
Q51 M AI How would you secure an MCP tool?
Tests: authorization. Answer: Treat MCP as an interface. Authenticate the caller, validate the token audience and tenant scope, validate tool arguments, apply least privilege and audit calls. A model’s choice is not authority to act. Confirm spec/SDK version because transport and auth evolved. Probe: Restrict get_policy to the actor’s tenant even if the model supplies another policy ID. Trap: Passing broad service credentials to an agent and relying on its prompt. MCP authorization; 2026 changes.
Q52 M AI What does a good golden set include?
Tests: evaluation design. Answer: Realistic anonymized user inputs, versioned source context, expert-approved expected behaviors, edge cases, adversarial and out-of-scope cases, plus slices by product, language and document quality. Keep a held-out set and adjudicate disagreement. Track score changes with model, prompt, retrieval and tool versions. Probe: A set of only easy positive cases rewards unsafe overconfidence. Trap: Reusing evaluation cases in prompt examples and calling the result generalization. Anthropic evals.
Q53 M AI Where do LLM calls go in Temporal?
Tests: durable orchestration. Answer: Put nondeterministic API calls in Activities. Workflow code emits replayable commands; Activities can retry, so external writes require idempotency. Set timeouts, retry policies and privacy-aware event history. Probe: Changing a running workflow’s command order can cause a nondeterminism error; version the code path. Trap: Calling the model directly from replayed workflow logic. Temporal.
Q54 S AI What does LangGraph persistence provide?
Tests: framework distinction. Answer: Checkpoints persist agent state at node boundaries for continuation and human interrupts. Design nodes so replay and side effects are safe; a checkpoint is not a substitute for a relational transaction around a policy issue. Probe: Compare checkpoint placement, tool idempotency and long-running business workflow needs to a Temporal design. Trap: Assuming every tool call happens exactly once. LangGraph JavaScript.
Q55 M AI How do you control latency and cost?
Tests: production economics. Answer: Track tokens, tool calls, retrieval size, model tier, cache hits and escalation per successful case. Use a cheaper deterministic gate where possible, cap context and iterations, and choose a model after eval by cost per accepted decision rather than cost per call. Probe: A cheaper model with more human corrections may cost more overall. Trap: Optimizing prompt tokens while false approvals rise. Gradion AI practice.
Q56 S AI How can AI speed engineering safely?
Tests: workflow discipline. Answer: Use it for exploration, draft tests and targeted refactors with a narrow spec, then inspect diffs, run checks and compare with product invariants. Track lead time, rework and escaped defects rather than generated lines. Share successful prompts or tools with context and failure modes. Probe: State one task you rejected because the model introduced a dangerous assumption. Trap: Equating faster typing with faster reliable delivery. Gradion careers.
Architecture and business Q57 to Q68
Q57 M Design What do you ask before drawing a system?
Tests: ambiguity handling. Answer: Identify user, journey, consistency/invariants, peak demand, data sensitivity, latency, recovery targets, integrations, budget and team ownership. Separate confirmed requirements from assumptions and ask which failures are acceptable. Then draw a minimal design and revisit at scale. Probe: For policy issuance, define what counts as an issued policy and when a user can safely retry. Trap: Starting with a queue and microservices before a user outcome. Gradion role.
Q58 M Design How do you estimate load honestly?
Tests: units and uncertainty. Answer: Convert users to operations and peak factor, multiply by fan-out and payload, then translate into concurrent work with latency. State assumptions as ranges and identify the capacity bottleneck. Benchmark the critical path. Probe: 100k commands/day is 1.16/s average, but a 20x peak is about 23/s before retries and fan-out. Trap: Sizing only from average MAU. AWS reliability.
Q59 M Design Strong versus eventual consistency?
Tests: boundary choice. Answer: Strong transaction boundaries protect invariants such as one accepted policy; downstream analytics, notifications and search can lag if the UI shows progress and retry semantics. Declare who owns the source of truth and how lag is measured. Probe: A notification failing should not roll back a committed policy, but a failed payment authorization may block acceptance. Trap: Calling all event-driven data “eventually consistent” without defining business behavior. PostgreSQL isolation.
Q60 M Design How do you decide build versus buy?
Tests: economic judgment. Answer: Compare differentiating need, integration fit, data/permission requirements, lifecycle cost, vendor exit and team capacity. Prototype the risky boundary and set a measurable decision date. Use managed identity or search where it meets the contract; own distinctive domain rules. Probe: A cheap SaaS quote can become costly when audit, tenancy and data residency needs appear. Trap: Counting only the initial license or sprint. Gradion consulting.
Q61 M Design How do you migrate a monolith safely?
Tests: evolutionary architecture. Answer: Identify bounded contexts and a measurable pain, isolate an interface, route one workflow through it, preserve data ownership and use compatibility tests. Migrate gradually with shadow reads or dual writes only when reconciliation is designed. Watch operational burden and stop if split cost exceeds benefit. Probe: Describe rollback when new service publishes an event the old monolith cannot read. Trap: Splitting tables by team without modeling transactions. Gradion modernization.
Q63 M Business Turn a vague client request into direction.
Tests: discovery. Answer: Ask whose job is painful, how often, cost of failure, present workaround and desired outcome. Map one journey and measurable acceptance criteria, then choose a small reversible experiment and align on ownership and rollout. Probe: “We need AI” might reduce to document classification with a rules baseline. Trap: Delivering a requested feature without establishing the problem. Gradion AI practice.
Q64 M Business What proves you owned a product lifecycle?
Tests: end-user contact. Answer: Name real users and the workflow you owned, initial signal, decision, release, support and feedback after adoption. Show one change made because users behaved differently than expected and measure its result or mark the metric unknown. Probe: For CoverGo, specify insurer/agent/end client personas and whether you personally saw feedback. Trap: Listing modules as proof of user outcomes. Gradion careers.
Q65 M Business How would you say no to a risky request?
Tests: judgment. Answer: State the goal you accept, explain the concrete risk and evidence, offer a smaller safe path with deadline, and record the decision. If a client wants an auto-approve agent without a labeled corpus, propose assistive review with a baseline and eval gate. Probe: Identify who has authority to accept residual risk. Trap: A flat refusal with no alternative. Gradion AI practice.
Q66 S Business Which metrics matter for a delivered feature?
Tests: outcome orientation. Answer: Pair adoption and task success with guardrails: quote completion, time to issue, rework/complaint rate, support burden and incident rate. Compare cohorts and a pre-launch baseline with careful attribution. Probe: Faster quote generation may be harmful if acceptance errors or abandonment increase. Trap: Counting shipped tickets as product impact. Gradion careers.
Q67 S Business What would you learn from the Shopware case?
Tests: source interpretation. Answer: Gradion publicly reports AI Co-pilot delivery and a cost outcome through an embedded team. The transferable lesson is to connect engineering choices to merchant workflows, adoption and economics. Do not infer Shopware’s private architecture or claim the cost change was caused by a specific model. Probe: Ask how the team measures merchant value and the cost to operate the feature. Trap: Repeating marketing figures as proof of a particular implementation. Case study.
Q68 M Business How do you present a failure?
Tests: accountability. Answer: State the original assumption, user impact and detection, what you personally did to mitigate, the root cause and a verifiable preventive change. Explain what you would decide differently with current evidence. Probe: If no metric was recorded, say so and describe what instrumentation you added. Trap: Turning a team failure into an unverified personal success story. Gradion careers.
Practical exercises
Attempt E01-E11 before reading the solutions. DSA examples are language-neutral pseudocode unless labeled otherwise. For each task, state input constraints and edge cases before implementation. Ask an interviewer whether Unicode normalization, overflow, stable ordering and mutation matter. Time targets assume a live session, not a hiring threshold.
Prompts
E01 DSA Two sum under duplicates 15 minutes
Return the first lexicographically ordered pair of indices (i,j), i<j, whose integer values sum to target, or none. Example [4,2,7,2], target 4 gives (1,3). Constraints: up to 1e6 items; negative values; duplicates. Hint: a map of earliest index by value. Consider integer overflow in a fixed-width language. Test empty, no pair and repeated equal values.
E02 DSA Longest unique substring 20 minutes
Return the length in Unicode code points of the longest contiguous substring without repeated code points. Example abca gives 3; empty gives 0. Constraints: up to 1e6 code points. Hint: sliding window and last-seen index. Clarify that code points differ from grapheme clusters and bytes. Test repeats, all unique and multibyte text.
E03 DSA Merge intervals 20 minutes
Given inclusive integer intervals [start,end], validate start<=end and merge overlaps, including shared endpoints. Example [1,3],[3,5],[8,9] gives [1,5],[8,9]. Constraints: unsorted input, up to 1e5 ranges. Hint: sort by start, scan. Test nested intervals, same endpoints and empty list.
E04 DSA Bounded worker queue 25 minutes
Design a queue with capacity K that supports offer, take, and clean shutdown for multiple producers/consumers. offer must return overload or respect a caller deadline; no silent unbounded buffer. Hint: use a bounded channel in Go or a bounded semaphore/queue abstraction in TypeScript. Test cancellation while full, shutdown while empty, and no lost accepted items.
E05 DSA Topological order 25 minutes
Given a list of jobs and dependency pairs (prerequisite,job), return a valid order or report a cycle. Example A->C, B->C permits A,B,C or B,A,C. Constraints: 1e5 vertices/edges. Hint: indegree, queue. Test isolated jobs, duplicate edges and cycle. Specify whether duplicate dependencies count once.
E06 DSA LRU cache 25 minutes
Implement get and put in expected O(1) with fixed capacity, updating recency on both successful get and put. Example capacity 2: put A, put B, get A, put C evicts B. Hint: map plus doubly linked list or ordered map with explicit behavior. Test capacity zero, overwrite and repeated eviction.
E07 SQL Idempotent policy issue 30 minutes
Design schema and transaction for POST /policies with (tenant_id, idempotency_key) and an insurer reference that must identify one active policy. Two replicas receive the same request concurrently. Return the original result for an identical replay, conflict for a different payload, and never commit two active policies. Include DDL, failure-after-commit scenario and an integration test outline. Hint: unique constraints and transaction, not a preflight SELECT alone.
E08 Frontend stale quote search 20 minutes
A React page fetches quotes as filters change. Network responses arrive out of order, and users see old tenant data briefly. Write a safe component design and test plan. Include request cancellation/stale-response guard, tenant-aware cache key, loading/error/empty states, and server authorization. Hint: browser cancellation is a performance aid; authorization remains on the server.
E09 Debugging rising p95 30 minutes
After scaling a Go API from 4 to 20 pods, p95 rises from an illustrative 180 ms to 2 s while CPU is 20%. Each pod has a PostgreSQL pool max of 30. Build a hypothesis tree, instrumentation checklist, immediate mitigation and experiment. Do not assume the pool is the only cause. Hint: 20 x 30 potential connections and pool/DB wait are starting clues.
E10 Workflow replay 20 minutes
A Temporal Workflow calls an LLM and an external issue-policy API inside workflow code, then a deploy reorders steps. Explain the failure and propose corrected orchestration, retry/idempotency boundary, versioning and replay test. Hint: Workflow replay compares commands; remote calls belong in Activities.
E11 AI evaluation design 30 minutes
Design a 100-case initial evaluation plan for an underwriting assistant using product documents and one read-only MCP policy tool. Define strata, label process, metrics, release gate, injection tests, human review and per-case trace. Every number is a proposed pilot assumption, not a claim about Tony’s implementation. Hint: separate retrieval, policy interpretation and final action.
Worked solutions
E01 solution
Scan left to right. For current value x at index j, seek target-x among earlier values; store the earliest index of each value. The first found j is minimal, but this does not always produce the lexicographically smallest pair globally: [5,4,1,0], target 5 gives (0,3) versus (1,2). To satisfy lexicographic order, scan all j, compare (i,j) pairs, and keep the minimum. Insert x only if absent. Expected O(n) time and O(n) space; use checked arithmetic for fixed-width types. Tests: [2,2] target 4 -> (0,1); no pair -> none; counterexample -> (0,3).
E02 solution
Pseudocode: maintain left=0, last={} and best=0. For each code point c at position right, if last[c]>=left, set left=last[c]+1. Set last[c]=right, update best=max(best,right-left+1). O(n) time and O(min(n,alphabet)) space. Iterate code points rather than UTF-8 bytes. Tests: abba -> 2; 你好你 -> 2; empty -> 0. If the product asks for visible characters, use grapheme segmentation instead and revisit the contract.
E03 solution
Reject malformed ranges; sort by start then end. Initialize output with the first range. For each next interval, if next.start<=last.end, set last.end=max(last.end,next.end), otherwise append it. O(n log n) time and O(n) result space (sorting may use additional memory). Tests: nested [1,10],[2,3] -> [1,10]; [1,1],[1,1] -> [1,1]. If touching but non-overlapping semantics are requested, change the inequality explicitly.
E04 solution
Go sketch (illustrative): a chan Job with buffer K, offer uses select { case jobs<-j: accepted; case <-ctx.Done(): timeout; case <-shutdown: closed }; workers select on jobs and cancellation. A coordinator closes the job channel only after producers stop, preventing send-on-closed panic, then waits for workers; define whether shutdown drains accepted work or abandons it. Capacity K bounds queued jobs, but workers add in-flight work, so total bound is K+workerCount. Test with barrier-synchronized producers, a full queue, and shutdown. Complexity O(1) per operation, O(K) queued memory.
E05 solution
Build adjacency lists and indegrees, deduplicating edges if inputs can repeat. Enqueue all zero-indegree jobs; pop, emit, decrement dependents and enqueue those reaching zero. If emitted count is smaller than vertex count, report a cycle. O(V+E) time and space. Tests: no edges emits every vertex; A->B,B->A reports cycle. For deterministic output, use a priority queue at O((V+E)log V).
E06 solution
Use a map from key to list node and a doubly linked list ordered most-recent to least-recent. get looks up and moves the node to the front; put updates/moves an existing node or inserts a new one, then removes tail if over capacity. Capacity zero stores nothing. Expected O(1) time per operation and O(K) space. In a multi-threaded server, wrap mutation with a lock or choose a concurrent design; algorithmic O(1) alone does not establish thread safety.
E07 solution
Illustrative PostgreSQL DDL: idempotency(tenant_id, key, request_hash, policy_id, response_json, PRIMARY KEY(tenant_id,key)); policies(id,tenant_id,insurer_ref,status,...) with a partial unique index on (tenant_id,insurer_ref) WHERE status='active'. In a transaction, claim the idempotency row (or lock an existing row), compare request hash for an existing key, create policy using a conditional state transition, store result and commit. Contending requests must wait or handle unique violations and then read the committed result. The row’s result cannot safely point to an uncommitted policy outside the same transaction. After a response loss, replay returns stored response. Test two concurrent connections and a crash/timeout after commit; verify exactly one active row. Index definition and status semantics require the actual product contract. PostgreSQL unique indexes.
E08 solution
Use a query key containing tenant and normalized filters. On change, cancel the previous request if possible and commit a response only if its request key still matches current state. Render an explicit loading state and keep prior tenant data hidden when tenant changes; show empty and error/retry separately. Server derives tenant from trusted identity and constrains SQL. Test deferred promises returning B before A, tenant switch while A is pending, 403 from server and keyboard navigation. React effects.
E09 solution
Potential capacity is 600 sessions after scaling versus 120 before. Check DB max connections, active/idle sessions, CPU/I/O/locks, pool wait, query durations, deployment timing and p95 trace breakdown. Mitigate by capping aggregate pool demand or reducing pod count if safe, stop runaway retries, and roll back if the change caused impact. Experiment with a fixed offered load and a smaller per-pod pool while measuring throughput and errors; profile slow SQL independently. A high pool maximum alone is a clue, not a diagnosis. PostgreSQL connection settings.
E10 solution
Workflow code must emit the same command sequence when replayed. Move LLM and issue-policy calls into Activities; assign the issue call a durable idempotency key and store result on the policy API. Keep retries bounded and distinguish permanent validation failures. Version running workflows with the SDK’s supported patching or worker-versioning strategy, then replay captured histories in tests before deployment. An Activity may execute more than once, so do not claim exactly-once effect. Temporal workflow definition.
E11 solution
Proposed 100-case pilot: 30 straightforward eligible, 20 clear ineligible, 20 ambiguous, 15 stale/conflicting policy versions, 10 low-quality/multilingual documents, 5 adversarial instructions. Have two domain reviewers label each with source passage, expected outcome and allowed abstention; adjudicate disagreement before scoring. Report retrieval recall at k, citation support, false approvals/denials, abstention, p95 latency, per-completed-case cost and tool authorization violations by slice. A release gate should have zero unauthorized tool access and no false approvals in a small critical sentinel set; do not claim statistical safety from 100 cases. Hold back independent cases and run shadow mode plus human approval before any automated decision. Anthropic evaluation design.
Worked system designs and business cases
These designs are authored interview scenarios inspired by the supplied CV and Gradion’s public domains. They do not describe Gradion’s or Tony’s actual internal systems. All volumes and targets in the designs are illustrative assumptions. Interviewers should be asked for real figures.
D01 Multi-tenant quote to policy platform
Actors and problem. Agents and brokers submit quotes across insurers; an end client accepts one; carrier integrations return decisions asynchronously. Support staff need a traceable lifecycle. Clarify insurer-specific rules, who can bind a policy, whether payment is required, whether one quote may be amended, and legal audit/retention with the client. Invariant: one acceptance command produces at most one active policy for a tenant/insurer reference, and unauthorized actors cannot see another tenant’s data.
Workload assumption. 200 tenants, 100,000 quote requests/day, a 20x peak over the daily mean, average payload 25 KB, and a 2 s user-visible quote response target for a synchronous carrier. That is roughly 1.16 requests/s average and 23/s peak before insurer fan-out. If each request fans out to three carriers, downstream peak is about 69 calls/s. Validate actual burst shape and carrier rate limits before sizing. At 2 s latency, about 46 peak quote requests could be in flight; pools, memory and in-flight carrier calls need more headroom than the daily average suggests.
API and model. POST /quotes uses a client-generated idempotency key; GET /quotes/{id} reports draft|submitted|quoted|declined|expired; POST /quotes/{id}/accept is a separately idempotent command; GET /policies/{id} exposes only authorized data. Model tenant, party, quote, quote_version, carrier_offer, acceptance, policy, idempotency_record and outbox_event. Use immutable quote/version identifiers so a late carrier response does not overwrite a newer request. A unique index enforces the policy invariant. Store price with currency and exact decimal representation; never use floating-point equality for money.
Flow. Authenticate actor, derive tenant, validate product/version, atomically record quote and command, then dispatch carrier calls with per-carrier deadlines and bounded concurrency. Return current state and a URL for progress if an insurer cannot answer within the interactive budget. Accept uses a transaction that checks version/state and writes acceptance, policy and outbox event under unique constraints. A publisher sends downstream events at least once; consumers dedupe. If an insurer’s bind API is an external side effect, coordinate with a durable workflow and insurer idempotency identifier, then reconcile ambiguous timeouts before marking the policy active. A browser can display pending state without promising issuance prematurely.
Reliability and security. Tenant scope applies in DB predicates, cache keys, logs and search. Authorize party/channel relationships at request and resource level. Timeouts and jittered retries honor carrier contracts; circuit-breaking or rate limiting isolates a sick carrier. Audit actor, rule version and decision reason while redacting sensitive documents. Instrument quote start-to-decision p95, bind success/duplicate prevention, carrier errors, pool wait, outbox age and tenant-isolation test failures. Use an SLO on the user journey and alert on error-budget burn, not only pod health.
Rollout and trade-off. Begin with a modular monolith and one transaction if one team owns the workflow; extract integration workers when carrier isolation or independent scaling merits it. Add schema fields compatibly, shadow old/new pricing for a subset, compare outcomes, then shift traffic per tenant or product. A queue improves resilience but adds lag and ordering complexity. At 10x volume, partition work by carrier/tenant and recheck DB locks, quotas and operational ownership. Interview follow-up: explain the failed response after commit and why the idempotency result must survive it. PostgreSQL isolation; AWS retry guidance.
D02 Underwriting AI assistant with human review
Actors and value. An underwriter needs policy-grounded evidence from changing product guidelines; a reviewer makes the final consequential decision. First establish the manual baseline: time per case, error and rework, document versioning, and how false approvals/denials affect customers. Test a narrow product and decision type. The assistant must never use a document from another tenant or invent a coverage rule. The CV reports an agent and eval harness but does not establish final production autonomy.
Workload assumption. A pilot has 500 cases/day with a 5x daytime peak, documents averaging ten pages, 100-case initial golden set and a manual review path. The average arrival rate is low; model latency, document parsing and review capacity dominate. Assume a 30 s assistant target and define which steps are asynchronous. Cost is model tokens plus embedding/index, storage, support and reviewer time per accepted case, not merely API calls.
Architecture. Ingestion validates file type and tenant, virus checks where appropriate, extracts page spans and metadata, versions the policy corpus, chunks by section, and indexes authorized text. Retrieval filters tenant/product/version before semantic ranking; it returns source references. A bounded orchestrator asks for missing facts and invokes read-only policy tools through server-side authorization. It produces structured proposed outcome, evidence spans, uncertainty and unresolved questions. A deterministic policy gate checks mandatory fields and routes uncertain/critical cases to human review. Store prompt/model/retrieval/tool versions, trace IDs and redacted artifacts for audit with appropriate retention.
Evaluation. Split retrieval quality from synthesis and action quality. Track coverage of authoritative passages, citation correctness, false approvals, false denials, abstention, tool permission violations, p95 latency, cost and human override by product/version. Test conflict between old and new guidelines, OCR errors and injected instructions. Use expert labels and held-out examples; compare to keyword search and a rules baseline. A green aggregate score is insufficient if a critical slice fails. Run offline eval, shadow mode, limited human-reviewed pilot and only then consider any bounded automation. The human keeps the final decision until evidence warrants a change.
Failure and rollout. If a new guideline is published, atomically mark its version and invalidate or re-index stale chunks; reject mixed-version evidence. If a tool times out, return an incomplete result and escalate. Treat retrieved text as data, block tool writes, and test adversarial PDF content. A model/version change triggers an eval run and canary. If the measured saving after reviewer corrections is negative, stop or narrow the feature. At 10x volume, batch ingestion, apply admission control to model calls and monitor queue age, while preserving tenant isolation. Anthropic evals; prompt injection; Temporal workflows.
D03 Commerce integration platform for a client
Provenance. Gradion’s public Shopware case reports an embedded AI product team and merchant features; its HomeToGo case reports 100+ third-party supply integrations. The design below is an authored hypothetical that combines the engineering concerns of multi-merchant commerce integrations. It is not a documented Shopware or HomeToGo architecture.
Actors and invariant. Merchants connect catalog/order systems, shoppers see current prices, and operators diagnose failures. Ask which platform owns inventory, how quickly price must refresh, how duplicate webhooks are identified, and how partial supplier outages affect checkout. Invariant: an order must not be charged twice, and a merchant must never see another merchant’s catalog or credentials.
Workload assumption. 1,000 merchants, 10,000 catalog updates/minute at peak (about 167/s), 100,000 shopper reads/minute (about 1,667/s), average update 2 KB. Change propagation within 60 s for most updates is a proposed target; checkout needs a stronger source-of-truth confirmation. Cardinality by merchant and supplier is skewed, so test hot merchants separately.
Model and flows. Use adapter interfaces with versioned canonical product/order schemas; keep raw source payload separately for forensic replay with retention limits. Webhook ingress authenticates/signature-checks the sender, records an idempotency key and event, then acknowledges quickly. A bounded queue drives validation, normalization, enrichment and upsert. Partition by merchant and item where order matters; dead-letter poison messages with a retry/replay tool. Read APIs serve an indexed materialized view or cache with freshness metadata. Checkout revalidates price/inventory with the authoritative supplier and uses a payment idempotency key; display a clear pending/failure state to the shopper.
Operations and trade-off. Measure ingestion lag, schema reject rate, per-supplier failure, item freshness, duplicate suppression, checkout success and p95 read latency. Limit per-merchant concurrency to prevent noisy neighbors. Deploy a new adapter behind shadow normalization and compare canonical records before promotion; keep a reversible mapping version. A universal connector abstraction can hide real supplier differences, so preserve explicit exceptions and contract tests. At 10x traffic, partition hot merchants, decouple read scaling, and check storage/index write amplification. Gradion integration description; HomeToGo case.
D04 Incident simulation and durable fix
Scenario. On Monday at 10:00 a release and scale-up coincide. Quote p95 moves from an illustrative 250 ms to 3 s, 5xx rises, carrier calls time out and users retry. The team initially suspects the new Go deployment. State what is known, what is suspected and what would falsify each hypothesis.
First 15 minutes. Declare user impact, freeze risky rollout, inspect error-budget burn and compare before/after traffic. Roll back or reduce concurrency if a safe mitigation exists. Trace the request: gateway wait, app CPU, DB pool acquire, SQL execution, lock wait and carrier span. Inspect pod count times pool size, active DB sessions, slow queries, retry volume and outbox age. Communicate a factual update with next checkpoint; keep a timeline. If retries are amplifying the outage, cap them and preserve idempotency. Avoid a DB restart merely because CPU is low.
Causal hypothesis and experiment. Suppose new replicas raise possible DB connections from 120 to 600 while a quote query also becomes slower on a skewed tenant. This is a hypothesis, not the revealed answer. Cap per-pod pool and concurrency, verify DB session/lock behavior, compare EXPLAIN (ANALYZE, BUFFERS) on representative data, and watch user journey p95 and error rate. If pool wait falls but query time remains, improve the specific query/index or restore the earlier plan. One change at a time makes attribution possible.
Aftercare. Add a capacity budget across replicas, alert on pool wait and queue age, integration-test duplicated acceptance, load-test a hot-tenant scenario, and add a staged rollout threshold tied to quote success. Document the actual contributing factors and why detection was late. The answer is strongest when the candidate distinguishes their real incident from this simulation. AWS reliability; PostgreSQL EXPLAIN.
Candidate story bank
These are prompts anchored to CV statements, not scripts that invent personal actions. Fill missing details with real facts before speaking. Use situation, task, your action, measured result, and what you learned. For every metric, know its denominator, measurement window and source.
| Story | CV evidence and suggested opening | Verify before interview | Follow-up likely |
|---|---|---|---|
| S01 CoverGo product ownership | Distribution platform; Party, Quote/Policy, Channel modules, multi-tenant foundations | Which insurer/agent/end users? What did you personally design and ship? What feedback caused a change? What production measure improved? | Cross-tenant test, release aftercare, stakeholder trade-off |
| S02 CoverGo AI assistant | Underwriting agent with Claude API, LangGraph, MCP, Qdrant/pgvector and eval harness | Exact workflow and tool permissions; sample cases, labels, baseline and false-decision rate; which parts were in production versus integration in progress | Prompt injection, human approval, policy version change, cost |
| S03 InvestaX trading integrity | CV reports OTC engine at 10,000+ daily transactions and 99.9% uptime; pool tuning | Measurement period and source; your code, p95/incident symptoms, before/after numbers, transaction invariant | Duplicate transaction, pool exhaustion, reconciliation |
| S04 InvestaX migration | Monolith to services, bounded contexts, gRPC/REST, SLIs/SLOs | One extracted boundary, how old/new coexisted, migration sequence and rollback; your decision authority | Why not remain modular, data ownership, service failure |
| S05 Select legacy modernization | Java/OSGi refactor, reusable SDKs, Python automation, real smart-home users | One concrete bug or security issue; reproduction, test, rollout, user impact and compatibility | Plugin lifecycle, backward compatibility, influence |
| S06 Hitachi ambiguity and delivery | GIS proof of concept; CV associates it with $2M+ funding; fleet platform scaling | Evidence of your contribution to funding decision; actual POC scope, client feedback and ownership; vehicle scale source | Hypothesis to demo to production, uncertainty in attribution |
| S07 Influence without authority | CV says standardized building blocks and team AI prompt patterns | Name a documented decision and dissent, whose mind changed, how evidence was gathered, and result | Mentoring seniors; handling a decision you lost |
| S08 Failure and learning | Choose a real incident across these roles | Initial mistake or missed signal, impact, mitigation, root cause, permanent fix and verification | What you would do differently; incident communication |
A cautious 60-second introduction
“I’m Tony, a hands-on full-stack engineer who has worked across insurance, fintech and IoT products. My strongest work is at the boundary of domain rules, data integrity and production operations. At CoverGo, I worked on the multi-tenant distribution platform across backend and UI modules, including quote/policy and channel workflows. Earlier, I worked on transaction integrity and service migration at InvestaX and on legacy and real-time systems at Select Technology. I’ve also built an underwriting AI assistant and an evaluation harness; I’m careful to separate the parts already running from integration still in progress. I’m interested in this role because it asks for an individual contributor who can code, diagnose difficult failures, make architecture calls and help teams carry those decisions into production.”
Revise this script to match exact personal contributions and years. If the live 7+ threshold arises, explain the actual dated employment record and correct the CV summary rather than debating the number. Do not describe the AI agent as autonomous production underwriting unless that is true.
Project deep-dive card
For each of S01-S04, prepare one page with: user and business pain; topology and data flow; invariant; volume and scale source; your specific decision and rejected alternative; one failure; security/tenant boundary; instrumentation; release/migration; measured result; what changed from feedback. Draw it from memory in five minutes. If a figure is unknown, label it unknown and describe how you would measure it now.
Questions for Gradion interviewers
Choose four to six for a real conversation. The interpretation notes are hypotheses, not judgments about Gradion.
| ID and audience | Natural question | Why and useful follow-up |
|---|---|---|
| I01 Recruiter | Which role posting and experience threshold applies to this process? | Supplied brief says 5+ while live posting says 7+; ask which team and evaluation criteria. |
| I02 Hiring manager | What does a strong first 90 days look like for a hands-on IC here? | Reveals ownership and success measures; ask which decision is most urgent. |
| I03 Client lead | Which client problem is currently hardest to turn into a technical direction? | Tests ambiguity and client exposure; ask how a decision is validated with users. |
| I04 Product lead | Which end users do engineers hear from directly, and how does feedback change releases? | Checks actual feedback loop; ask for a recent example. |
| I05 Tech lead | Where do architecture decisions live, and how are they revisited? | Tests influence and documentation; ask who can challenge a decision. |
| I06 Tech lead | Which workload most often sets your reliability limits: data, external APIs, UI or operations? | Opens a system discussion without assuming their stack; ask for a recent trade-off. |
| I07 Tech lead | How do teams balance a client’s near-term deadline against debt or a reliability risk? | Reveals how “say no” works; ask for a decision they changed. |
| I08 AI lead | What evidence is required to move an AI pilot into a user-facing product? | Tests eval, safety and economics; ask who labels failures and owns review. |
| I09 AI lead | How are tool permissions, tenant data and human approval enforced for agents? | Tests actual boundaries without presuming MCP; ask about adversarial tests. |
| I10 Peer engineer | What is the most painful production incident you learned from recently? | Reveals learning culture; ask what changed afterward. |
| I11 Peer engineer | How do you make code review valuable when AI drafts much of the code? | Tests ownership and tests; ask for a review that caught a domain error. |
| I12 Delivery lead | How are distributed teams and clients kept aligned on contracts and releases? | Relevant to Gradion’s multi-location work; ask which handoff fails most often. |
| I13 Tech lead | Which languages or frameworks would this team actually use in the first project? | The JD is deliberately stack-agnostic; ask for versions and why chosen. |
| I14 Hiring manager | Where has the team decided not to use AI, and why? | Tests judgment and ROI; ask what baseline won instead. |
| I15 Product lead | How do you measure merchant or end-user value beyond delivery velocity? | Links to public commerce cases; ask which metric moved after a release. |
| I16 Recruiter | What are the actual interview stages, languages and expected preparation? | Avoids inferring a proprietary process; ask about live codebase or take-home. |
Two-week study plan
Assumption: 14 days x 90 minutes = 21 hours. Each session spends 35 minutes learning, 30 minutes answering or building without notes, 15 minutes reviewing errors, and 10 minutes adding evidence to a story card. If the interview is sooner, do days 1-3, 5, 8, 10, 12 and 14 first; the complete handbook remains the reference. If daily time is lower, keep the same priority order and reduce optional breadth.
| Day | Focus and linked material | Completion signal |
|---|---|---|
| 1 | Diagnostic; S01; Q01-Q08 | Explain one request end-to-end in 4 minutes; score weak areas. |
| 2 | Go internals; Q09-Q13; E01-E02 | Solve two DSA tasks and explain slice aliasing, race and profile choice. |
| 3 | Node/TS; Q14-Q18; E04 | Draw event loop versus worker/DB waits; bound a worker queue. |
| 4 | DSA E03, E05, E06 | Solve with edge tests and complexity before viewing answers. |
| 5 | Data integrity; Q19-Q22; E07 | Defend one durable idempotency transaction and concurrent test. |
| 6 | Distributed systems; Q23-Q30; D01 | Walk duplicate event, timeout and tenant boundary. |
| 7 | Frontend; Q31-Q38; E08 | Show safe stale-response and accessible failure states. |
| 8 | Production; Q39-Q46; E09; S03 | Build incident timeline, mitigation and two hypotheses. |
| 9 | Release/migration; D04; S04-S05 | Tell a migration with a rollback and one measured result. |
| 10 | AI mechanisms; Q47-Q53; E10 | Explain retrieval, tool auth, prompt injection and replay. |
| 11 | AI economics/eval; Q54-Q56; E11; S02 | Present eval slices and production boundary honestly. |
| 12 | System design; Q57-Q62; D02-D03 | Complete a 35-minute design with units and trade-offs. |
| 13 | Client/product; Q63-Q68; S06-S08; I01-I16 | Tell feedback and disagreement stories; choose 5 questions. |
| 14 | Two 30-minute mocks; final review | Score 0-4 in correctness, reasoning, judgment and communication; revisit scores <3. |
Mock rubric
For each dimension, 0 = absent/incorrect; 1 = fragmentary; 2 = workable with gaps; 3 = correct with trade-offs; 4 = precise under follow-up. Dimensions: correctness, causal reasoning, implementation/design judgment, and communication. Record a quoted claim or behavior that justifies each score, one focused improvement, and a retest question. These are practice signals, not hiring predictions. To run a live mock later, ask one question, answer it aloud or in text, receive one probe, then see the model answer and score.
Last-day review sheet
- Mechanisms: request -> transport -> runtime -> transaction -> downstream -> browser; CPU time versus waiting; slice aliasing; Go race/context; Node event loop; TypeScript runtime validation; React effects and stale requests.
- Invariants: tenant scope, unique policy, idempotency after lost response, transaction boundary, outbox duplicate, insurer timeout reconciliation.
- Production: p95 with trace and pool/DB evidence; SLI on user journey; mitigate before tuning; compatible schema rollout; rollback and verification.
- AI: rule/search baseline; versioned authorized retrieval; false approval versus review cost; held-out eval by slice; tool auth in code; hostile documents; human gate; Temporal Activity side effects.
- Stories: S01 user feedback, S02 AI production boundary and eval, S03 measured incident, S04 migration, S07 disagreement, S08 failure. Verify every number and your contribution.
- Ask: I01 threshold discrepancy, I02 first 90 days, I08 pilot gate, I13 team stack, I14 decision not to use AI.
Research notes and source index
Research checked on 4 October 2026. The supplied brief is the role authority for this package; the live posting is a separate, possibly newer variant. The CV is candidate-provided evidence. No actual interview rounds, client assignment, deployed versions, production agent autonomy, independent performance benchmarks or privileged Gradion architecture were verified. Company case figures are self-reported in first-party materials. All technical questions are curated practice, not claimed historical interview questions. For the interview, confirm language/runtime and database versions before giving version-specific implementation details.
| Domain | Primary source and use |
|---|---|
| Role and business | Gradion live role, careers, full-stack services, AI practice |
| Documented case studies | Shopware, HomeToGo, Swiss credit redesign |
| Computing and Go | HTTP overview, Go memory model, Go race detector, Go diagnostics, net/http |
| Node, TypeScript, React and browser | Node event loop and worker pool, TypeScript handbook, React effects, React memo, Web Vitals, WAI forms |
| Data and reliability | PostgreSQL isolation, EXPLAIN, index usage, AWS retries, Google SRE SLO example, OpenTelemetry |
| AI and workflows | Anthropic eval design, prompt-injection guidance, tool use, Temporal workflow definition, LangGraph JS, MCP July 2026 changes |