The benchmark pairs a Modern Standard Arabic (MSA) baseline track with an Iraqi track across six axes: dialect comprehension, dialect generation, bidirectional MSA-Iraqi translation, Iraq-specific knowledge, official-document field extraction, and safety. It was built from 340 originally authored, dually reviewed items, with statistically audited answer positions and Wilson intervals on every published score. A pilot evaluation of 27 systems found the MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders. Official-document extraction confined every system to 32-56. Two dedicated Arabic models scored below a size-matched generalist on the Iraqi track. The safety-hardened tier of the newest model family deterministically refused innocuous dialect-comprehension items as policy violations.