FlowCraft Systems · Research preview

The AI writes the code.
Judgment is the job.

DevSense drops you into a simulated engineering environment with an AI agent, a production incident, and a codebase you've never seen — then watches how you investigate, delegate, and decide. We measure how you think, not how fast you type.

// investigate  ·  delegate  ·  evaluate  ·  decide  ·  verify

“You are entering a simulated engineering environment. Work with your AI engineering agent as you would on a real team. Investigate, delegate, make decisions and solve the situations that emerge.”

The briefing every candidate receives
The shift

Assessment is measuring the wrong thing.

Traditional coding tests ask whether you can personally produce a correct implementation under artificial constraints. That is precisely the part of the job that is being automated.

01

Implementation is becoming cheap

Capable agents now generate most of the code. Being the fastest typist in the room stopped being a differentiator somewhere around last year.

02

The scarce skill moved upstream

Framing the real problem. Knowing what evidence you need. Calibrating how much to trust the agent. Spotting the change that is locally correct and systemically dangerous.

03

Nobody is testing for it

The interview loop has not caught up. DevSense makes higher-order judgment observable by putting it under load — inside an environment that behaves like real engineering work.

The simulation

A fictional company. A real engineering loop.

You join one company for the whole session and work through several connected episodes — an ambiguous bug, a production incident, an agent-written change that needs reviewing. Context accumulates. So does the consequence of your decisions.

Sense

Read the code, the diffs, the logs, the telemetry, the git history, the tickets, the incident record.

Reason

Form hypotheses. Prioritise them. Go looking for the evidence that would actually discriminate between them.

Delegate

Direct the agent. Decide how much autonomy the task, the risk and the reversibility deserve.

Verify

Treat the agent's output as evidence, not truth. Check that the outcome you wanted actually happened.

Adapt

Abandon the dead end. Recover from the bad call. Recognise when the evidence is finally sufficient.

You direct the agent — you don't edit the code

Every change to the system goes through the AI engineering agent. Code is evidence you reason about, not the medium you produce work in. You can navigate it, inspect diffs, read tests, reject changes and demand verification.

The agent is competent, and sometimes wrong

It has variable autonomy and variable reliability. It may do excellent work, overreach, or confidently recommend something unsafe. Every flaw that counts is discoverable from evidence in the environment. We measure judgment, not mind-reading.

Failure is a state transition, not a game over

A poor decision degrades system health, breaks something, or creates customer impact — and then hands you the chance to notice and recover. How you recover is often better evidence than the original mistake.

The session adapts to you

Once we have sufficient evidence of a competency, we stop probing it and spend the remaining time elsewhere. Strong candidates are not punished with endlessly harder problems. 30–60 minutes, breadth first.

What we measure

Seven dimensions of engineering judgment.

Every conclusion traces back to something you actually did — a diff you inspected, a hypothesis you dropped when the evidence turned, a deployment you refused to approve.

01

Problem framing

Establishing the actual problem, the constraints, the unknowns and what success would look like.

02

Reasoning & hypothesis quality

Building plausible explanations, seeking discriminating evidence, and updating when it contradicts you.

03

Judgment under uncertainty

Knowing when to investigate, when to act, when to wait, and when not to do the irreversible thing.

04

Resourcefulness

Finding the route to information — telemetry, git history, ADRs, incident records, service ownership.

05

Critical evaluation

Treating agent output as evidence. Separating assertion from proof. Catching the confident wrong answer.

06

Risk & systems thinking

Blast radius, reversibility, downstream dependencies, compatibility, data integrity, second-order effects.

07

Adaptive execution

Recognising diminishing returns, changing strategy, recovering from mistakes, verifying the outcome.

Honest claims

What we will and won't say about you.

Assessment credibility is a product feature. We would rather make a narrow claim we can defend than a broad one we cannot.

We do claim

  • Demonstrated performance across defined engineering-judgment competencies, within the simulator.
  • Specific, inspectable behavioural evidence for every conclusion we reach.
  • Developmental feedback you can act on — strengths, gaps, and the decisions that showed them.

We do not claim

  • That a score means someone is objectively a better engineer.
  • That any rating predicts on-the-job performance. That requires longitudinal validation we have not done.
  • That a single number captures it. The model underneath is multidimensional, and stays that way.
Early access

We're looking for people to break this.

DevSense is in early validation. We are running sessions with engineers and students across a wide range of experience, and having senior engineers independently review the traces against our automated assessment.

If you want to take a session, review one, or talk about using this with your team or your campus — tell us which.

Button not opening your mail app?

no account · no leaderboard yet · 30–60 minutes