AGENTIC AI HACKATHON 2026 Product Space

The model proposes evidence.
The code validates it.

HireFlow reads a job description and a stack of resumes, then produces a screening report where every claim points back to the exact line of the source document it came from. When a claim can't be grounded, HireFlow refuses it and asks the question instead.

0 match scores shown anywhere
5 of 50 claims refused by the verifier
91s to screen a 5-candidate pool

The idea in one sentence

The LLM never gets the final say on its own evidence.

The model returns candidate claims with quotes. A verifier written in plain Python — no model call anywhere inside it — then goes and locates each quote in the source document. If the quote can't be found, the claim is refused and the requirement is flagged for a human instead of being marked satisfied.

That single boundary is what makes the output defensible. The component that decides truth structurally cannot hallucinate, because it cannot generate text.

ingest Job description
+ resumes
extract Model proposes
claims + quotes
verify Pure-code verifier
locates every quote
map Requirement status
+ evidence matrix
ask Interview question
for what stayed open

Requirement statuses

Assigned by explicit rule, not left to the model's discretion.

StatusRule
Met Explicit evidence satisfies the requirement
Partial Explicit evidence satisfies part of the requirement
Unverified Related evidence exists, but the required detail is not demonstrated
Absent No relevant evidence found at all

Unverified is the status that matters. A resume saying "containerised microservices with Docker" against a requirement of "Kubernetes in production" is not a match — it's a question to ask.

Model integration

Claude does the reading. Python decides what's true.

The LLM is used for exactly one job: reading a resume and a job description, then proposing candidate claims with the quotes that support them. Every one of those proposals is handed to a verifier written in plain Python, with no model call anywhere inside it, which goes and locates the quote in the source document. A claim the verifier cannot ground is not downgraded or hedged — it is refused, and the requirement stays open for a human.

Because the trust boundary sits in code rather than in the model, the agent stays useful on a weaker model. Swapping the model never changes what HireFlow is willing to assert, which is the property we actually care about.

Model-agnostic by design

The pipeline talks to any OpenAI-compatible /chat/completions endpoint through HIREFLOW_BASE_URL and HIREFLOW_MODEL. Moving to a different provider is a configuration change, never a rewrite.

Batched for throughput

One call per candidate, run concurrently, instead of one per requirement. That is the difference between roughly 19 minutes and about 91 seconds for a five-candidate pool.

Structured output, handled

JSON is enforced by system prompt, tolerant parsing and a single repair retry rather than a vendor flag — so the client does not depend on a provider-specific guarantee.

Nothing trusted from the browser

API routes hold the prompts and the model key server-side. The browser never receives credentials, and the recorded run renders with no model access at all.

Real output, one candidate

Every finding carries the quote it came from.

These are unedited findings from the recorded run — Senior Backend Engineer at Northwind Commerce, candidate C-01. Notice the difference between the two Spring Boot findings: the same quote, read twice, at two different levels of confidence.

Met REQ-06 · Has Docker experience

The capstone project explicitly used Docker and Docker Compose to run the application locally, demonstrating Docker experience.

"Ran the application locally with Docker Compose."

Projects — Campus Event System · aman-verma.md
Unverified REQ-03 · Production experience with Spring Boot

Spring Boot is used in a university capstone project, but the resume does not demonstrate production experience with Spring Boot.

"Built a campus event management system using Spring Boot and Docker."

Projects — Campus Event System · aman-verma.md
Became an interview question

"Your resume shows Spring Boot use in a university capstone project. What experience have you had deploying and maintaining Spring Boot applications in a production environment, and what responsibilities did you personally own after deployment?"

Absent REQ-09 · Kafka or event streaming

The resume contains no evidence of Kafka or another event-streaming platform.

Refused by the verifier

The model proposed related text. The verifier could not locate a supporting quote, so the claim was refused and this requirement stayed open for a human.

Partial REQ-01 · 5+ years backend experience

The resume evidences six months of professional software engineering internship experience, but it does not show at least five years of professional backend development.

"Fixed bugs in a Java monolith and wrote unit tests."

Experience — TechNest · aman-verma.md

The product

A working screening workspace, not a mockup.

Intake, requirement coverage, the candidate pool, the evidence matrix and a full audit trail — every screen reads from a recorded run, so the whole app works with no model access at all.

A verified verdict showing a candidate's evidence checked against a requirement
Verified verdict Every claim carries its own provenance — finding id, confidence, the model that produced it, and a timestamp.
Evidence matrix mapping candidates against job requirements
Evidence matrix Candidates against requirements, so unverified and absent gaps are visible at a glance rather than buried in prose.
Audit trail showing which information produced each insight
Audit trail Which information produced each insight — no claim in the report exists without a trace back to the source line.

The recorded run

Numbers read off the running app, not estimated.

10 requirements parsed from one job description 6 must-have · 4 nice-to-have
5 candidates screened synthetic fixtures, one deliberately ambiguous
50 evidence checks run one per candidate-requirement pair
45 quotes located in the source document every one matched back to its document
5 claims refused by the verifier each became an interview question
21 requirements left open for a human the honest answer, not a guess

How the 50 findings actually landed

  • 20 met
  • 9 partial
  • 16 unverified
  • 5 absent

Stack & architecture

A LangGraph pipeline with one real conditional edge.

evidence_sufficient? routes either to a verified finding or to a generated interview question. The verifier node is plain Python and is covered by unit tests — the part that decides truth is the part with no model in it.

Web

Next.js App Router with React and Tailwind. API routes hold the prompts and the model key server-side; the browser never sees credentials.

Agent pipeline

Python and LangGraph against an OpenAI-compatible chat endpoint. Batched one call per candidate, run concurrently.

The verifier

Pure Python, no model call. Tries to locate each proposed quote in the source document, and refuses the claim when it can't.

Audit trail

Findings carry their own provenance — finding id, confidence, the model that produced them and a timestamp — visible behind Why? on every claim.

Two constraints that shaped the design

  • No structured-output guarantee. JSON is enforced by system prompt, tolerant parsing and a single repair retry — not by a vendor flag. This kept the client provider-agnostic.
  • Slow calls. One call per requirement would have taken about 19 minutes for five candidates. Batching one call per candidate and running them concurrently took ~91 seconds.

Built for the Agentic AI Hackathon 2026

12 of 13 required capabilities covered. Capability 5 — explicit candidate grouping — was deliberately deprioritised: with five candidates, the evidence matrix and the natural-language query already covered it.

founder@hire-flow.dev