Start with the language reality
MENA workflows may combine Modern Standard Arabic, dialects, French, English, transliteration, and specialized vocabulary inside a single task. Evaluation should first describe who communicates, which documents are used, and where language switching occurs.
- Map languages by user and workflow
- Include dialect and transliteration where relevant
- Preserve domain terminology in the test design
Measure the task, not the model in isolation
Retrieval quality, groundedness, citation fidelity, extraction accuracy, and appropriate refusal can matter more than a generic score. Criteria should correspond to the decision or service the system supports.
Build governed evaluation material
Representative examples need provenance, access rules, versioning, and review. A small, well-governed set that reflects actual edge cases can be more informative than a large public benchmark disconnected from the institution.
- Record provenance and permitted use
- Version examples and expected behavior
- Review coverage as workflows and language evolve
Keep accountable people in the loop
For high-consequence workflows, evaluation must include escalation and human review—not only automated scoring. The aim is to understand where the system assists reliably, where it must defer, and how an operator can inspect its output.