Ask HN: Anyone interested in building a harness-only benchmark?

4 pointsposted 9 hours ago
by GodelNumbering

Item id: 49181014

5 Comments

theChris-in

9 hours ago

Interesting idea. If you can manage the infra, I can put together a replicable test suite.

GodelNumbering

8 hours ago

I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal

theChris-in

8 hours ago

We can do a mix of general use (as in user stories) plus a few academic benchmarks.

So you have any specific ideas?

You can hmu at iam@thechris.in

GodelNumbering

8 hours ago

Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_...

> We can do a mix of general use (as in user stories) plus a few academic benchmarks.

Yup sounds about right. Generally speaking, higher the distinct contributors, more likely it is to capture the distribution of real-life usefulness.

> So you have any specific ideas?

Only that the problems that get picked should be easy to evaluate in isolation and should test the harness capability rather than model's knowledge/capability.

theChris-in

8 hours ago

> should test the harness capability rather than model's knowledge/capability.

Then we need a provenance for model inference, generalized. This should be interesting. We would be trying to deterministically generalize a baseline "can do this" for models... Maybe categorize by parameter class.