Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative…
ABSTRACT A continuing challenge in aerospace materials is the search for alloys that have desired functional properties to operate at higher temperatures with lower densities. This improves aero‐engine efficiency and reduces CO 2 and other harmful emissions, aligning with aviatio…