Implementation is becoming cheap
Capable agents now generate most of the code. Being the fastest typist in the room stopped being a differentiator somewhere around last year.
DevSense drops you into a simulated engineering environment with an AI agent, a production incident, and a codebase you've never seen — then watches how you investigate, delegate, and decide. We measure how you think, not how fast you type.
“You are entering a simulated engineering environment. Work with your AI engineering agent as you would on a real team. Investigate, delegate, make decisions and solve the situations that emerge.”
Traditional coding tests ask whether you can personally produce a correct implementation under artificial constraints. That is precisely the part of the job that is being automated.
Capable agents now generate most of the code. Being the fastest typist in the room stopped being a differentiator somewhere around last year.
Framing the real problem. Knowing what evidence you need. Calibrating how much to trust the agent. Spotting the change that is locally correct and systemically dangerous.
The interview loop has not caught up. DevSense makes higher-order judgment observable by putting it under load — inside an environment that behaves like real engineering work.
You join one company for the whole session and work through several connected episodes — an ambiguous bug, a production incident, an agent-written change that needs reviewing. Context accumulates. So does the consequence of your decisions.
Read the code, the diffs, the logs, the telemetry, the git history, the tickets, the incident record.
Form hypotheses. Prioritise them. Go looking for the evidence that would actually discriminate between them.
Direct the agent. Decide how much autonomy the task, the risk and the reversibility deserve.
Treat the agent's output as evidence, not truth. Check that the outcome you wanted actually happened.
Abandon the dead end. Recover from the bad call. Recognise when the evidence is finally sufficient.
Every change to the system goes through the AI engineering agent. Code is evidence you reason about, not the medium you produce work in. You can navigate it, inspect diffs, read tests, reject changes and demand verification.
It has variable autonomy and variable reliability. It may do excellent work, overreach, or confidently recommend something unsafe. Every flaw that counts is discoverable from evidence in the environment. We measure judgment, not mind-reading.
A poor decision degrades system health, breaks something, or creates customer impact — and then hands you the chance to notice and recover. How you recover is often better evidence than the original mistake.
Once we have sufficient evidence of a competency, we stop probing it and spend the remaining time elsewhere. Strong candidates are not punished with endlessly harder problems. 30–60 minutes, breadth first.
Every conclusion traces back to something you actually did — a diff you inspected, a hypothesis you dropped when the evidence turned, a deployment you refused to approve.
Establishing the actual problem, the constraints, the unknowns and what success would look like.
Building plausible explanations, seeking discriminating evidence, and updating when it contradicts you.
Knowing when to investigate, when to act, when to wait, and when not to do the irreversible thing.
Finding the route to information — telemetry, git history, ADRs, incident records, service ownership.
Treating agent output as evidence. Separating assertion from proof. Catching the confident wrong answer.
Blast radius, reversibility, downstream dependencies, compatibility, data integrity, second-order effects.
Recognising diminishing returns, changing strategy, recovering from mistakes, verifying the outcome.
Assessment credibility is a product feature. We would rather make a narrow claim we can defend than a broad one we cannot.
DevSense is in early validation. We are running sessions with engineers and students across a wide range of experience, and having senior engineers independently review the traces against our automated assessment.
If you want to take a session, review one, or talk about using this with your team or your campus — tell us which.
Button not opening your mail app?