Responsible AI · Model Performance Baseline

Scored consistently.
Grounded in evidence.
Verified monthly.

Every interview is checked for whether it asks relevant questions and probes vague answers further. Every resulting score is checked for consistency and groundedness.

99.5%Justification groundedness
0.07Score SD (overall)
99.2%Question relevancy
93.6%Depth follow-up rate

Computed on live production interviews. Recomputed monthly.

Methodology

How every interview is verified.

Two agents. Four checks. Each operating independently so the interview process is auditable end to end.

01Question GenerationAvya generates role-specific questions from the capability mapRelevancy verified
02Adaptive ProbingThin or vague answers trigger follow-up questions in real timeDepth verified
03ScoringEach competency scored independently on a 0–5 scale with written justificationConsistency verified
04Justification AuditAn independent judge model traces every justification to the transcriptGroundedness verified

All checks run on production interviews, not synthetic test data. The evaluation agent and questioning agent operate as independent systems.

Stage 01 · Evaluation Agent

Scoring Reliability & Groundedness

Whether written score justifications are accurate, evidence-based, and whether repeated scoring yields stable results.

Justification Groundedness
99.5%pass rate

Justifications traced to candidate statements by an independent judge model. 0.1% hallucination rate — flagged and excluded.

Scoring Consistency
0.07pts SD · overall

Test–retest standard deviation on the 0–5 scale. Per-competency: 0.13 pts. Scores are stable across repeated evaluations.

What this means

If two evaluators independently score the same interview, their scores differ by less than 0.15 points on a 5-point scale. Every justification points to something the candidate actually said.

Why it matters for hiring

When a hiring manager asks 'why did this candidate score 3.8?', the justification is verifiable against the transcript. No black-box scoring. Every number has a paper trail.

Stage 02 · Questioning Agent

Relevance & Adaptive Depth

Whether generated questions test the intended competency and whether ambiguous responses are probed further.

Question Relevancy
99.2%on-target

Every question independently judged to directly test the stated competency — not a tangential or unrelated area.

Adaptive Depth
93.6%follow-up rate

Thin, vague, or incomplete candidate responses trigger a deeper follow-up rather than being accepted as-is.

What this means

When Avya is configured to assess 'distributed systems design', it asks about distributed systems — not tangential topics. And when a candidate gives a surface-level answer, Avya probes deeper 94% of the time.

Why it matters for hiring

A candidate who memorizes a textbook definition gets probed until they demonstrate real understanding — or reveal the gap. Surface-level answers don't pass as competence.

All metrics at a glance

Full summary with definitions and results.

AgentMetricDefinitionResult
EvaluationJustification Groundedness% of justifications traceable to candidate statements99.5% pass
EvaluationScoring ConsistencyStandard deviation on test–retest (0–5 scale)0.07 overall
QuestioningQuestion Relevancy% of questions that directly test stated competency99.2% on-target
QuestioningAdaptive Depth% of vague answers triggering follow-up probes93.6% follow-up

Honest limits

A report containing only favourable findings would be the less credible one.

📐

Process soundness, not prediction

These four checks confirm the evaluation process is reliable. Downstream hiring outcomes require separate validation on live populations.

🔬

A precision floor exists

Consistency SD of 0.07 is tight but not zero. Very small scoring differences below the detection threshold require more data to resolve.

📅

A point-in-time baseline

Models and prompts evolve. A result valid today is not a permanent certificate. Baselines are recomputed whenever models or prompts change.

🔗

Scope is internal to Avya

These metrics measure Avya's interview engine in isolation. They do not account for upstream screening or downstream panels.

Performance baselines are recomputed monthly. Metrics are re-run whenever models or prompts change, on expanding candidate and scenario sets. Findings are published openly — including when they regress.

Hear It From Them

Why Global Enterprises trust Zeko AI!

Global standards verified by Fortune 500 Companies

  • “Zeko AI helped us evaluate the gap between a good hire and the right hire along with the calibre of people we needed: so we keep bringing in people most companies struggle to reach."

    Talent Acquisition Leader

    Enterprise Technology Company · 2,500+ employees

  • "Getting every hiring manager to hire the same way is usually the hardest part. With ZekoAI, that adoption came easy - and hiring got simpler across the board."

    HR Business Partner

    IT Services Company · 10,000+ employees

  • "In logistics, we hire fast and in huge numbers. Zeko AI let us manage it all in one place - without ever letting quality slip."

    Recruitment Lead

    Logistics Company · 5,000+ employees

  • "Hiring for a project used to take months for us. With Zeko AI it's weeks now — and that speed, without losing quality, is what changed how we deliver."

    Delivery Head

    IT Services Company · 5,000+ employees

  • "Zeko AI ensures every candidate is evaluated against the same benchmarks - every time. Better decisions, less bias, stronger hires."

    Vasundhara Chakravarthy

    Pierian Services

Trusted by Global Enterprises

Workflows curated by sector, configured by role.

Persistent Systems
Infosys
Schneider Electric
Nagarro
Shadow Fax
Pierian Services
Orion Innovation
Recro
Coditas
Persistent Systems
Infosys
Schneider Electric
Nagarro
Shadow Fax
Pierian Services
Orion Innovation
Recro
Coditas

Enterprise-grade security with Responsible AI built-in

SOC2
ISO27001
GDPR
SOC2
ISO27001
GDPR

Act Now

Build a Consistent, Audit-Ready Hiring Process

Standardize interviews across geographies and improve hiring quality with Zeko's AI platform.

  • Trusted by 150+ enterprises

  • SOC2 · GDPR · ISO27001

  • 4.8/5 Average Candidate Rating

Act Now

Build a Consistent, Audit-Ready Hiring Process

Standardize interviews across geographies and improve hiring quality with Zeko's AI platform.

  • Trusted by 150+ enterprises

  • SOC2 · GDPR · ISO27001

  • 4.8/5 Average Candidate Rating

Act Now

Build a Consistent, Audit-Ready Hiring Process

Standardize interviews across geographies and improve hiring quality with Zeko's AI platform.

  • Trusted by 150+ enterprises

  • SOC2 · GDPR · ISO27001

  • 4.8/5 Average Candidate Rating