HeyAbrar
All selected work

HR technology

Three evaluation layers. No black-box hiring score.

I engineered the structured evaluation pipeline behind a candidate-screening platform that combined video analysis, adaptive assessment, and interview intelligence without turning model output into an unexplained hiring decision.

Client
Confidential enterprise HR technology company
My role
AI systems and full-stack engineer
Platform
AI web platform
3evaluation layers
80%screening workload reduced
94%reasoning vocabulary consistency

Technology

Multimodal language modelsOpenAI WhisperAzure Speech SDKOpenCVLangChainWebRTCFastAPIPostgreSQLReactNext.js

Every score needed evidence a recruiter could challenge

High-volume technical hiring created inconsistent CV reviews and expensive phone screens before interviewers had reliable evidence of a candidate's knowledge or communication.

The legal and product requirement was explicit: every automated score needed criterion-level reasoning that a recruiter could inspect rather than a single opaque ranking.

Made model output reviewable

I owned work across the structured evaluation pipeline, audio and video quality controls, adaptive-exam model, and recruiter-facing intelligence brief.

The design treated model output as reviewable evidence. Scores, reasons, confidence, and source quality travelled together through the system.

Architecture

Evidence travelled with every score

Schemas, confidence, source quality, and criterion-level reasoning formed one evaluation record that could be compared, audited, and tested by a human interviewer.

  1. 01

    Constraint

    Free-form model explanations used inconsistent criterion names and could not support reliable audits.

    Decision

    Constrain every evaluation to a versioned JSON schema with an allowed criterion vocabulary, score, evidence, and confidence.

    Impact

    Reasoning outputs used the defined vocabulary correctly in 94% of validation cases and became comparable across roles.

  2. 02

    Constraint

    Fixed technical tests spent too many questions confirming what the system already knew about a candidate.

    Decision

    Use item-response theory to select each next question from the current ability estimate and recalibrate question parameters from production data.

    Impact

    Recalibrated question banks produced 18% lower score variance and targeted knowledge boundaries more precisely.

  3. 03

    Constraint

    Browser recording quality varied enough to distort transcription and delivery analysis.

    Decision

    Run a local pre-submission signal-quality check and carry transcription confidence into the final evaluation.

    Impact

    Candidates who rerecorded after a warning produced audio with 22% better speech-recognition error rates.

Outcome

Faster screening with a visible reasoning trail

  • The highest-volume screening step required approximately 80% less recruiter effort.
  • A 200-candidate pool could be processed overnight.
  • Interviewers using the intelligence brief probed predicted knowledge gaps 67% more often.
  • Every evaluation retained explicit reasoning and data-quality indicators for human review.

Explainability improved the product as well as the audit trail. Requiring evidence for every score exposed weak criteria early and gave interviewers something concrete to test.

Abrar

Start a project

Building an AI decision-support product people must trust?

I can help design evaluation schemas, evidence trails, human review, and quality controls around uncertain model output.

Discuss your AI platform