smithclay
3 hours ago
More benchmarks comparing effectiveness of various AI SREs is welcome and overdue: especially vendor-vs-vendor comparisons. One area that I think is going to be really interesting and important is the best way to emulate complex IT environments for evals.
Some related work I recommend checking out: - https://arxiv.org/abs/2609.33023 (new last week!) - https://github.com/SREGym/SREGym - https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-... (Clickstack's version) - https://github.com/grafana/o11y-bench (Grafana's version)
yildizfatih
2 hours ago
thanks for the link. Can you provide an example for "complex IT environments", we would love to explore it.