cafkafk
7 hours ago
Author here. This came out of reading Denmark's national biological threat assessment (first update since 2020), which splits its AI finding in two: not much help for people without a professional background, meaningful help for people who already have one. I suspected that it wouldn't reach as far on its own, so I've tried to get some of its ideas across, as well as some of my own.
I think it's important now that some frontier models approach to making biology guardrails is near a blanket ban. The literature doesn't necessarily support that approach, and I think it's a second order consequence of how models are trained to be unbiased, rather than e.g. taking more opinionated stances.
Another point is that if automation actually displaces people in pharma and biotech research, those are significantly more likely threat actors, meaning that managing the social consequences of large language models for critical sectors is directly a part of ensuring they don't have actual "doomsday" like outcomes.
Most alignment discussions I've seen focus on cybersecurity, and neglects e.g. the reality of counterterrorism, social impacts that lead to bad outcomes, and so on.
My hope is we can see frontier labs adopt a more sophisticated approach to guardrails, and change the discourse to be more holistic and grounded in the real risks.