The model proposes evidence. The code validates it.
HireFlow reads a job description and a stack of resumes, then produces a screening report
where every claim points back to the exact line of the source document it came from. When a
claim can't be grounded, HireFlow refuses it and asks the question instead.
The LLM never gets the final say on its own evidence.
The model returns candidate claims with quotes. A verifier written in plain Python — no
model call anywhere inside it — then goes and locates each quote in the source document.
If the quote can't be found, the claim is refused and the requirement is flagged for a
human instead of being marked satisfied.
That single boundary is what makes the output defensible. The component that decides truth
structurally cannot hallucinate, because it cannot generate text.
ingestJob description + resumes
extractModel proposes claims + quotes
verifyPure-code verifier locates every quote
mapRequirement status + evidence matrix
askInterview question for what stayed open
Requirement statuses
Assigned by explicit rule, not left to the model's discretion.
Status
Rule
Met
Explicit evidence satisfies the requirement
Partial
Explicit evidence satisfies part of the requirement
Unverified
Related evidence exists, but the required detail is not demonstrated
Absent
No relevant evidence found at all
Unverified is the status that matters. A resume saying
"containerised microservices with Docker" against a requirement of
"Kubernetes in production" is not a match — it's a question to ask.
Model integration
Claude does the reading. Python decides what's true.
The LLM is used for exactly one job: reading a resume and a job description, then
proposing candidate claims with the quotes that support them. Every one of those proposals
is handed to a verifier written in plain Python, with no model call anywhere inside it,
which goes and locates the quote in the source document. A claim the verifier cannot ground
is not downgraded or hedged — it is refused, and the requirement stays open for a human.
Because the trust boundary sits in code rather than in the model, the agent stays useful
on a weaker model. Swapping the model never changes what HireFlow is willing to assert,
which is the property we actually care about.
Model-agnostic by design
The pipeline talks to any OpenAI-compatible /chat/completions endpoint
through HIREFLOW_BASE_URL and HIREFLOW_MODEL. Moving to a
different provider is a configuration change, never a rewrite.
Batched for throughput
One call per candidate, run concurrently, instead of one per requirement. That is the
difference between roughly 19 minutes and about 91 seconds for a five-candidate pool.
Structured output, handled
JSON is enforced by system prompt, tolerant parsing and a single repair retry rather
than a vendor flag — so the client does not depend on a provider-specific guarantee.
Nothing trusted from the browser
API routes hold the prompts and the model key server-side. The browser never receives
credentials, and the recorded run renders with no model access at all.
Real output, one candidate
Every finding carries the quote it came from.
These are unedited findings from the recorded run — Senior Backend Engineer at Northwind
Commerce, candidate C-01. Notice the difference between the two Spring Boot findings: the
same quote, read twice, at two different levels of confidence.
MetREQ-06 · Has Docker experience
The capstone project explicitly used Docker and Docker Compose to run the application
locally, demonstrating Docker experience.
"Ran the application locally with Docker Compose."
UnverifiedREQ-03 · Production experience with Spring Boot
Spring Boot is used in a university capstone project, but the resume does not
demonstrate production experience with Spring Boot.
"Built a campus event management system using Spring Boot and Docker."
Became an interview question
"Your resume shows Spring Boot use in a university capstone project. What experience
have you had deploying and maintaining Spring Boot applications in a production
environment, and what responsibilities did you personally own after deployment?"
AbsentREQ-09 · Kafka or event streaming
The resume contains no evidence of Kafka or another event-streaming platform.
Refused by the verifier
The model proposed related text. The verifier could not locate a supporting quote, so
the claim was refused and this requirement stayed open for a human.
PartialREQ-01 · 5+ years backend experience
The resume evidences six months of professional software engineering internship
experience, but it does not show at least five years of professional backend
development.
"Fixed bugs in a Java monolith and wrote unit tests."
The product
A working screening workspace, not a mockup.
Intake, requirement coverage, the candidate pool, the evidence matrix and a full audit
trail — every screen reads from a recorded run, so the whole app works with no model
access at all.
Verified verdict
Every claim carries its own provenance — finding id, confidence, the model that
produced it, and a timestamp.
Evidence matrix
Candidates against requirements, so unverified and absent gaps are visible at a glance
rather than buried in prose.
Audit trail
Which information produced each insight — no claim in the report exists without a
trace back to the source line.
The recorded run
Numbers read off the running app, not estimated.
10requirements parsed from one job description6 must-have · 4 nice-to-have
5candidates screenedsynthetic fixtures, one deliberately ambiguous
50evidence checks runone per candidate-requirement pair
45quotes located in the source documentevery one matched back to its document
5claims refused by the verifiereach became an interview question
21requirements left open for a humanthe honest answer, not a guess
How the 50 findings actually landed
20 met
9 partial
16 unverified
5 absent
Stack & architecture
A LangGraph pipeline with one real conditional edge.
evidence_sufficient? routes either to a verified finding or to a generated
interview question. The verifier node is plain Python and is covered by unit tests — the
part that decides truth is the part with no model in it.
Web
Next.js App Router with React and Tailwind. API routes hold the prompts and the model
key server-side; the browser never sees credentials.
Agent pipeline
Python and LangGraph against an OpenAI-compatible chat endpoint. Batched one call per
candidate, run concurrently.
The verifier
Pure Python, no model call. Tries to locate each proposed quote in the source
document, and refuses the claim when it can't.
Audit trail
Findings carry their own provenance — finding id, confidence, the model that produced
them and a timestamp — visible behind Why? on every claim.
Two constraints that shaped the design
No structured-output guarantee. JSON is enforced by system prompt,
tolerant parsing and a single repair retry — not by a vendor flag. This kept the client
provider-agnostic.
Slow calls. One call per requirement would have taken about 19 minutes
for five candidates. Batching one call per candidate and running them concurrently took
~91 seconds.
Built for the Agentic AI Hackathon 2026
12 of 13 required capabilities covered. Capability 5 — explicit candidate grouping — was
deliberately deprioritised: with five candidates, the evidence matrix and the
natural-language query already covered it.