AI researcher Andrej Karpathy, who joined Anthropic earlier this year, recently put Claude Opus 5 through a unique coding test, demonstrating how benchmarks for large language models (LLMs) are ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results