Evaluating a model you cannot read
When the output is prose, exact-match metrics tell you almost nothing. A useful eval set starts from the decisions the text has to support.
Start from the decision
Write down what a reader must be able to do after reading the output. Then grade that, not the wording.