Select Page

Humanity’s Last Exam

Benchmarks are interesting.

Here’s the deep thought – at what point in the overall benchmark process will AI inject bias into the benchmark test?  And to what end?  Maybe not so deep a thought.

Humanity’s Last Exam has been bantered about extensively.   Here’s a great place to catch up on it: Humanity’s Last Exam

My musings: check out the crazy difficulty of the questions:

So 2500 questions of this caliber of difficulty.  The top AI models hit 20% accuracy in answering. 

 

 

 

I would also note the Calibration Error, which is affirms that “ Given low performance on Humanity’s Last Exam, models should be calibrated, recognizing their uncertainty rather than confidently provide incorrect answers, indicative of confabulation/hallucination. To measure calibration, we prompt models to provide both an answer and their confidence from 0% to 100%. ”  The better performing models – OpenAI o3 and o4-mini and Gemini 2.5 Pro – also have better Calbration Error numbers.

AI Action plan, and stuff

AI Action plan Here ya go folks - this is the current administration's AI Action Plan:  https://www.ai.gov/action-plan Here are some words from the current administration about preventing "woke AI" in the federal government...

read more

UltraEdit

I first used UltraEdit sometime in the late 90s.  I loved it back then. When I went to work for TQuist in 2003 I purchased another copy. I just now retired my last XP system, and so I've decided to get the latest copy.  I hope it's aged well. They've changed...

read more