josefcub
9 hours ago
You know, I took this when there were only a handful of people on the leaderboard. Things I didn't know:
20-40 minutes is 2 hours if you're like me and want to make sure you're actually thorough answering questions. Aisa (the agent) was great at asking pertinent questions and keeping the conversation somehow both on-topic and on-going. It was not nearly as laborious as I thought it would be, despite how long it took.
I scored 89. I'm still in the top 30 on the leaderboard. I have no idea how I scored that with as little experience at work as I have with AI. At the time AISA first debuted here, we hadn't even gotten access to agentic models at my employer, just chatbots. I learned literally everything at home with local models.
I'm not going to argue a high score, but that feeling of disbelief is still real. I think this is a great product for evaluating your personal skills according to a fairly broad but well-defined rubric. The fact that you get to read the analysis itself and why you were graded as you were is great for self-improvement.
If I had an ask, it would be to actually download that analysis and the original conversation log. It was comprehensive enough that I'd love to turn it into actionable steps for improvement without having to copy-paste so extensively.
Thanks for making an evaluation system that works so well!
Ozzie-D
an hour ago
Thank you for sharing your experience. The calibration has "settled" a bit more since then and you might score perhaps a bit lower (or not) but might be worth a retry especially if you have learned new things.
For example we have a fatigue model (that tracks how bored the user is and adjusts) a retrain model (that retrains on the frontier knowledge in AI everyweek) a golden-shadow model (that looks at what you did NOT say to make a prediction about how good you might be on HOW you say) and a learning-agility model. These are all new.
It's not perfect but it's improving with every assessment and I have a ton of product backlog notes to make it even more advanced.