miellaby
5 days ago
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
ivo-42
5 days ago
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
wuschel
4 days ago
I had the pleasure to speak to your former COO - so glad your organisation exists and publishes its amazing work!
brcmthrowaway
5 days ago
How do you cleanse the data at this scale?
ivo-42
5 days ago
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
stephantul
5 days ago
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
tfburns
2 days ago
Tom here. Thanks for the note! I think we also interacted via X on this topic? Certainly interested to see what the comparison looks like.
xvfLJfx9
5 days ago
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
ivo-42
5 days ago
Stay tuned, this is just the beginning. Merging with Cohere will help in scaling, too.
idiotsecant
5 days ago
How exactly would an open project do that?
soundworlds
5 days ago
Awesome work. The work truly needs it!
doctorpangloss
5 days ago
How much should German authors be paid, $3,000 per book, like Anthropic paid?
Why expose yourself to this liability?
gewetensleegte
3 days ago
Anthropic was fined. It's a rather big difference.
zelphirkalt
5 days ago
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
davidjfelix
5 days ago
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
gewetensleegte
3 days ago
> the US industry
to be a bit more exact
slow_typist
5 days ago
How, is the training data public?
WhyNotHugo
5 days ago
This was really pleasant to read — especially so for the presentation of a new model.
I felt like a learnt a lot of details about the entire process. Questions arose during my reading, and searching for the answers led to more learning.
No marketing BS, lots of actual information. And their level of openness is really neat!
p-e-w
5 days ago
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
ofjcihen
5 days ago
Hopefully this becomes the new standard.
It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
Loquebantur
5 days ago
Be cautious what you wish for. Tools don't tell you what to do with them.
Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.
Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.
Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.
adrianN
5 days ago
Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
zanderwohl
5 days ago
The select cabal who believe in an eschaton they're actively trying to bring about, as well.
skinfaxi
5 days ago
Would you apply this reasoning to the proliferation of nuclear weapons?
edit: why is the parent rationale sensible for AI and not nuclear technology?
orbital-decay
5 days ago
Absurd and incomparable.
skinfaxi
5 days ago
Why? That seems like a shallow dismissal. Why is a world changing technology okay in the hands of a small cabal in the one case and not another?
orbital-decay
5 days ago
"Why is the wheel not okay to gatekeep but the nuclear bomb is?" Even the framing is manipulative from the beginning. By asking this question you're already assuming they're in any way comparable. They are not even remotely equivalent and the entire comparison is utterly absurd.
skinfaxi
5 days ago
AI is more like a nuke than a wheel. Do you disagree? Can you suggest a less manipulative framing? I am personally a proponent of open source AI and models but I found this cabal framing strange when we do indeed rely on this kind of control for other world-altering technologies. And AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
adrianN
4 days ago
AI is more like the general purpose computer than a nuke. It can be used to greatly accelerate research and development of all kinds of things. Of course the war on general purpose computing has been going on for many years now, but so far freedom still has a couple of strongholds left.
orbital-decay
4 days ago
AI is more like the wheel (can be used for anything) than the nuke (purposeful weapon of indiscriminate mass destruction), and there's more than two categories, so there's no need to prove proverbial Godwin's law by appealing to extremes. But I still can answer the ridiculous question about incomparable things with as much good faith as I have for it: gatekeeping a generic tech in advance based on ridiculous hypotheticals is ridiculous. That's a charitable interpretation, there's also uncharitable one: that it's nonsensical drivel made up by malevolent cranks that benefit from this strawman being used in public discussions.
>AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.
The wheel already enabled almost every tech. Without the wheel nothing you see around yourself would be possible.
Forgeties79
5 days ago
Dude ChatGPT is not the equivalent of a device that can level a major city killing millions in a flash. It is self evident. This entire discussion is ridiculous.
Nuclear weapons are a wholly unique threat to mankind.
Loquebantur
5 days ago
That's a misconception and manipulative framing on your part.
While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
AI allows anybody to enact essentially anything. And your "level a major city" is just a small task really. The problem there is your lack of imagination, not the actual impossibility of that task.
komali2
5 days ago
> AI allows anybody to enact essentially anything
No it doesn't. Where is your evidence for this?
> While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.
This doesn't exist. Why are you pretending it does?
stale2002
5 days ago
So, in other words, none of this is about anything related to the actual technology and instead people are talking about the made up, fake technology that you read in Sci-fi book.
Thats the annoying part about these conversations. People try to smuggle in the conclusion of "And now I wave a magic wand that does literally anything, by magic" when discussing stuff that everyone can use and see right now that clearly isn't that.
Forgeties79
5 days ago
It’s not manipulative at all. This is a comparison of nuclear weapons and LLM’s. That is manipulative.
komali2
5 days ago
Proliferation of data doesn't mean proliferation of the thing itself.
There's a massive barrier between knowing how to build a nuke and actually building it.
Also, isn't proliferation the basis for MAD? Well if we believe in that, naturally it means that the world would be safest if every individual had their own nuke ;)
taneq
4 days ago
Nuclear weapons aren’t that good an analogy, AI is closer to biological weapons.
ofjcihen
5 days ago
Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.
But the alternative just seems… so much worse to me?
A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.
And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.
Loquebantur
5 days ago
Yes, a "select" (rather, self-selected) elite "controlling" AI according to their wishes, what could go wrong?
Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.
When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.
"General" intelligence isn't present in humans, why does it need to be in AI?
throwaway27448
5 days ago
> When access to AI gives you abilities you cannot use responsibly
I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.
skinfaxi
5 days ago
> When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that.
What do you mean "cannot"? As in you are granted abilities that have no responsible use?
Loquebantur
5 days ago
"Cannot" as in presently cannot. That includes the case of things that have no responsible use.
You live in a curated world and rarely or never encounter such things. Precisely because your environment is curated that way.
Look at how you can't buy WMDs. They have no responsible use for you.
Zigurd
5 days ago
The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.
But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.
As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?
GolDDranks
5 days ago
I find it odd that more people don't realize that we are talking about risk of elevated *general intellectual-domain capabilities* as a resource. To be clear, I'm not claiming that LLM + RF is necessarily THE technology that poses the risk, the risk is in recursive self-improvement and whatever technologies will result.
It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.
Zigurd
5 days ago
The reason I use coding and vulnerabilities as an example is that it is a concrete example. It is what people pay for now when they buy AI. And the risks are specific and can be examined in detail. Some threads on this board currently show that even these more concrete and specific risks are often overblown, with LLMs finding low risk bugs and sucking up resources to evaluate and fix them.
Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.
I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?
Loquebantur
5 days ago
What poses the risk is the combination of abilities past a certain point enabling you to do basically anything.
While being unable to judge whether you should in the first place.
vincnetas
5 days ago
ai cant move atoms at unlimited rate and also have limited energy. so your claim that "basically anything" is a bit of a stretch.
Loquebantur
5 days ago
That's what you believe, but you might be wrong.
nradov
5 days ago
So you're saying that AI might invent infinite energy and other technologies indistinguishable from magic? I mean I can't prove that's impossible but you're making a pretty lame statement that essentially amounts to a religious prophecy.
Loquebantur
5 days ago
You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:
What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.
-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.
Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.
Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.
It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.
Zigurd
5 days ago
This is analogous to the problem of spam, which is a problem about five minutes younger than email. Before spam you had to buy ads in the back pages of magazines you think target vulnerable demographics. Meta already spews fraudulent ads in horrific volume.
In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.
Loquebantur
5 days ago
You compare to laughably irrelevant things why?
"Ambitious frauds" haven't "explored all the angles".
You imply "LLMs" to be and stay less intelligent than humans, in particular yourself. You're mistaken.
fwn
5 days ago
Whenever classic p(doom) sentiments are explained through a text with obvious LLM markers, I wonder whether I am looking at a superhuman persuasion attempt.
..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.
On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.
The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.
Maybe through some terminally online overconsumption, like with social media? I wouldn't know.
With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.
It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s
jemmyw
5 days ago
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.
Not arguing that we shouldn't strive to do far far better here, better is by our own imagining. There is no development scale. There are no aliens or prior non human civilizations to compare against. So we're not underdeveloped. We are as we are. For all we know we're at peak capacity and humanity will never be better.
computerdork
5 days ago
Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.
Wonder what safeguards Kolibri uses to prevent this? Or if they even can
zanderwohl
5 days ago
I don't think that AI is as much of a boost to bioweapons as people think. Lab work doesn't get easier just because the experiment design part does.
Loquebantur
5 days ago
You can simulate things.
Given enough smarts, you can simulate anything. Quantum AI is an active goal.
komali2
5 days ago
Mate throughout this thread you've been escalating LLMs to be basically the equivalent of a magic wand.
I think we should stay in reality for now. They write good rust and bad emails.
lawandjustice
4 days ago
They also get gold at IMO for 15 dollars and solve Erdos problems.
computerdork
4 days ago
eh, they are also used extensively in medical research using alpha fold, and in math, and in legal for one thing. And this is of course only the beginning, and IMHO, it'd take a lack of foresight to not start putting layers and layers of safeguards now.
Yeah, to me the hugging face hacking is a harbinger of what's to come, not just in the world of tech but eventually to other fields too.
komali2
3 days ago
If I described a computer to you in 1998 I'd have said "they are also used extensively in medical research using alpha fold, and in math, and in legal for one thing." Maybe not alpha fold, idk if that's some new thing. But you get my point? They're a very useful tool, sometimes it feels like magic what we can do with them... But it's not magic. That's why I said magic wand, the OP was implying LLMs can (or soon can) manipulate matter at an atomic level. That's orders of magnitude removed from where we are and I disagree that there's any evidence that LLMs will inevitably reach that point in our lifetimes.
computerdork
3 days ago
Ah, I see, LLM's aren't able to do incredible things and you were just commenting on that. Sounds reasonable
Zigurd
5 days ago
Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.
UberFly
5 days ago
Enter LLMs to help through all those pesky hard parts.
shawabawa3
5 days ago
The hard part is probably finding the lab equipment and chemicals without being noticed, and choosing not to use it to manufacture drugs instead which would be much more profitable
throwaway27448
5 days ago
Intelligence was never the bottleneck tho
happosai
5 days ago
For terrorists it is tho. Four lions is basically a documentary. The only clever and creative terror attack happened 25 years ago.
throwaway27448
5 days ago
What's stopping them from using claude today? What could anyone possibly do to stop them from using open models? This line of thought seems like corporate/political/pr pandering more than a meaningful concern.
skinfaxi
5 days ago
Didn't Mexican cartels kidnap telco workers to build them separate infra? We expect terrorists to be less resourceful?
happosai
4 days ago
Well yes. drug cartels have a revenue source.
Miscreants with even half a braincell will find themselves in illegal drugs business over terrorism.
Political and religious fanatics with even a half a braincell figure out terrorism will only damage their cause.
People who end up as terrorists have already failed the two above tests.
throwaway27448
5 days ago
I think the unethical things are happening already with boutique firms. I admit I don't get the concern.
OtomotO
5 days ago
> Humanity has the serious problem of being underdeveloped in the "spiritual" department.
As is shown to us by filthy rich people every day.
Or did you mean the burglar in the fawellas?
98753579909754
5 days ago
[dead]
dang
5 days ago
The paper mentioned is here: https://tej.as/blog/aleph-alpha-kolibri.
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)
api
5 days ago
His point about regulation and innovation is great and I wish more people thought like that.
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
dang
5 days ago
(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)
amelius
5 days ago
I wish universities would take it upon themselves to curate the training sets for these models.
user
5 days ago
mm321
4 days ago
Very interesting. I wonder, if there are plans to add French in a future version, or will Kolibri remain bilingual? (English/German). The domains include industry/aerospace, so there would be some demand for French, presumably.
giancarlostoro
5 days ago
Isnt OLMO basically open like this? My understanding is people have recreated the model from the same data sources with repeatable responses, or reasonably close to the original (since LLMs never answer the same).
gunalx
5 days ago
LLMs can be configured to basically be deterministic. It's only really bad for performance because language is not deterministic.
kingcauchy
5 days ago
Yeah the pdf alone is awesome as a learning tool.
zwaps
5 days ago
Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.
Bravo team!