Paseer AI Center

Evaluation and quality assurance

Arabic and English test sets for each capability, measuring attribution, hallucination, bias, cost and time, red-team tests, and deployment gates that safely disable the capability when a threshold fails. ENT-AI-040..045 G-05.4 NFR-REQ-003

Adding a test set is not available in this prototype: how it is added and who approves it are not yet decided

Published thresholds

AI-DATA-006ENT-AI-042AI-SEC-004

Latest evaluation results

2026-09-24 02:00 · Groups AR/EN
CapabilityCases AR / ENUnsourcedHallucinationBiasPrompt injectionGate

Red team

Red Team · Prompt injection · permission tests ENT-AI-044
  • Prompt injection inside a PDF attachment PDF212 patterns · 0 successes
    Held
  • Paragraph-level privilege escalation80 scenarios · 0 leaks
    Held
  • Re-identification of masked individuals2 of 150 re-identified AI-SEC-006
    Needs handling

Release gate · Canary

Extract commitments v2
  • Full evaluation passed
    16 Sep · all thresholds
  • Canary to 10% for 72 hours
    Started 17 Sep · 0 violations so far
  • 50%
    Automatically if metrics stay within thresholds
  • 100% · Capability owner approval
    Automatic safe disable on any critical failure