Mold Linker Version 3.0.0 Release – Rewritten in Rust

245 pointsposted 2 days ago
by roflcopter69

168 Comments

lrvick

2 days ago

The linux distribution I co-maintain uses mold as our bootstrap linker to bootstrap rust itself, and it saved us -hours- on long version-by-version build chains. Mold being in c meant we could build it very early and use it as the default linker distro wide and enjoy build speedups everywhere.

Now sadly we will have to fork and maintain the c version as mold2 forever.

Rust is not actually the right tool for all problems.

jdc-pub

a day ago

Like some other commenters, I do not understand why using a cached binary isn’t sufficient. Why not use the mold3 binary and build that first, and then cache it and never rebuild it again? I’m guessing it has to do with hermiticity and build provenance guarantees in your distro.

Do you have any reading material that you can share that might help me understand better? Docs for the distro, or an issue tracker I can search through?

lrvick

a day ago

https://stagex.tools

In short, the entire distro is always built in one-shot at any given commit, and we can only rely on cached binaries from a past release if they or nothing in their supply chains changed. Given rustc depends on almost everything, mold3 would be built far too late to be useful for the most expensive build in the whole tree, which is rust.

Recursive dependencies would break our threat model, so we cannot use any rust tools to bootstrap rust. We bootstrap rust from llvm which we bootstrap from gcc which we bootstrap from tinycc which we bootstrap from M2Planet and so on back to 186 bytes of machine code.

Anyone getting a rust package from stagex must be able to build the entire tree up until that package and get the same hash, with no binary dependencies, thus removing any trust in maintainers.

rererereferred

16 hours ago

I wonder if a rust-to-c tool would help you, your mold2 would just be the cached result of that. Would that C code still be treated as a binary dependency since it would not be very human readable?

How about a rust-to-wasm process, you store the wasm binary, but have a human readable wasm interpreter that only implements enough for mold to work? Would that be trusty enough?

lrvick

14 hours ago

If a rust-to-c tool could fully handle this automatically, and said tool only depended on an early language like go or c, that would be very useful for many problems including this one.

nh2

a day ago

What _is_ your threat model?

If you can boostrap one version of stagex (so that its mold is trustworthy), I don't see why you can't use that to build another version of stagex.

As you say, you can rely on the cached binary because nothing in the old version's supply chain would change; it is pinned, and as long as you can build that, everything is fine.

lrvick

a day ago

The threat model is trust no single person or computer.

In order for you to not have to trust us, you must clone our repo, of only source code, and build from zero to our released binary hashes. The shorter we can make the time for that to be possible, the more people we can convince to do it and ideally sign and publish their matching hashes. The more people we convince to do it, the less risk of us as maintainers being able add a backdoor without anyone noticing.

Mold was a tool to shave hours off the time most people have to spend doing a from-scratch verification. Forcing them to build a whole tree to get to rust to et to mold3, to then use that to build the whole tree a second time, would directly work against the goal of minimizing full tree verification time.

dwattttt

a day ago

> In order for you to not have to trust us, you must clone our repo, of only source code

> rust from llvm which we bootstrap from gcc which we bootstrap from tinycc which we bootstrap from M2Planet and so on back to 186 bytes of machine code.

While the effort is laudable, the fact that these are built from source is not enabling me to verify them. Rust, LLVM, gcc, tinycc, all the way back to that 186 seed are prohibitively large to verify.

I'm not even in a position to verify the delta to these projects over a single day. I'm having to trust _someone_, many someones.

lrvick

a day ago

An operating system distribution owns specifically the compiling and distribution links of the supply chain. It is our job to prove we produce artifacts accurately and faithfully from upstream source code without giving ourselves any ability to tamper with the sources or the resulting artifacts.

Our specific responsibility is to limit the number of people users have to trust to only the actual authors of the software they are installing.

sunshowers

a day ago

This is entirely self-inflicted on your part!

> Rust is not actually the right tool for all problems.

Rust is certainly the right tool for this problem, your own decisions notwithstanding.

lrvick

a day ago

Self inflicted because we want to protect as many people as possible from supply chain attacks. That is a very real engineering problem we have to solve that extremely high risk end users of our work depend on.

Most popular Linux distros take a position of hoping and praying supply chain attacks do not target them. I am not convinced this will go well for them in the post AI world, but hey, I also hope I am wrong.

sunshowers

14 hours ago

There are plenty of other points in the design space here. For example, I think once a sufficient number of people have built a binary and verified that it is the same hash (publishing the hashes to an irrevocable transparency log), caching those binaries doesn't increase the supply chain surface. The point you have chosen is quite out there, and it isn't reasonable to expect others to conform to it, and it certainly isn't reasonable to say that Rust isn't the right choice for a project because it negatively impacts you.

lrvick

13 hours ago

> caching those binaries doesn't increase the supply chain surface.

Okay, but what if all the signers are people that a new user does not know? Who is to say all those signatures are not a bunch of made up AI identities?

This is why we make it trivial for anyone to clone our tree and build from zero and get the same result at any release with no required binaries of any kind anyone has to verify the provenance of. This is called full source bootstrapping. The cheaper we make that, the more people that will do it and the higher the chances users will see a signature from someone they personally trust.

hitekker

12 hours ago

Your arguments sound reasonable to me. It’s a bit worrying to hear people playing make-believe about security.

I’ve been told “no one thinks a programming language always makes software safer” but the comments here suggest many actually do believe just that. The problem you’ve defined is a threat to the “my language is a silver bullet” belief.

sunshowers

12 hours ago

I think you might be fundamentally misunderstanding distro maintainers' role in the ecosystem. Your job is not to question your upstreams' development practices or technologies used, beyond the basics like it being FOSS -- it is to adapt to them.

lrvick

6 hours ago

Upstreams naturally know a lot less about supply chain security than we do as that is our core expertise. We leave deciding what features are correct for their users to them, and handle supply chain integrity.

Often having to make lots of small patches to preserve upstream functionality while fixing their obvious security and determinism bugs. Or we have to ignore autogenned code and figure out bootstrapping ourselves. Many upstreams just put binaries in their source code and call it reproducible.

For instance, XZ published a malicious hand-packed archive of their code that many distros use because it has all the auto-generated code already. We totally ignored that archive and pulled the code that was actually reviewed, and ran autogen ourselves. In doing so we were never impacted by the XZ attack.

We are obligated to do anything we can to protect our users, even if most thing doing so is paranoid.

karavelov

a day ago

Why not use LLD for linking Rust? You already must have the LLVM for Rust to be buildable, so that do not add any dependency.

lrvick

a day ago

We use the llvm linker, lld, to link mold at the earliest point we can, then use mold exclusively after that. We are an LLVM native distro so we build llvm right away as the system compiler used to build the whole tree instead of gcc.

LLD would work fine for the whole tree, and did previously, but is much much slower than mold which is why we switched to mold.

Having to build all dependencies of rust including python and openssl and everything else before being able to use mold erases most of the full-tree build speedups as the path to rust is already about 80% of the full tree build time.

We are a distro that mandates independently verified 100% deterministic builds from source for any given release commit of the tree, so we have to build the whole tree several times for every release.

aseipp

a day ago

Can't you just keep using Mold2 for the initial toolchain bootstrap and then build mold3 with Rust as part of the "final" set of toolchains that are to be used to compile everything beyond that? It isn't perfect but it's congruent with how many other components in most open reproducible bootstrapping flows work; ie using GCC 4.4 or whatever just to bootstrap GCC 10. It's annoying but it only needs to be done once and then you just use it to bootstrap newer components forever. Doesn't help with compile times, though.

lrvick

a day ago

I would agree if not for the fact rust depends on almost the entire tree, so doing what you propose would increase our full tree verification time from 6ish hours to 12ish hours erasing any wins mold was giving us in the first place.

dwattttt

a day ago

I may be misinterpreting the parents point, but you're not building E.g. tinycc for it to be a final artefact; why not continue to build mold2 in its current place in the build chain? It's not going anywhere, much like tinycc isn't.

lrvick

a day ago

We have to rebuild tinycc any time a dependency of tinycc changes. Likewise we would have to rebuild mold every time a dependency of mold changes, including rust and all dependencies of rust. It is a problem of almost full tree dependency recursion.

dwattttt

a day ago

Yes. I'm saying that mold3 gets added to the end of the build tree, and mold2 remains where it is currently, which doesn't change how often you need to build those and let's you continue to use mold2 to speed up the builds (which I've now seen is what you'll likely do).

My point though was that you already build many end-of-life projects just so they can be used in the build, mold2 can now one of those rather than being built for its use in the final distro.

lrvick

14 hours ago

I indicated as much in my original comments. Pinning mold2 and backporting relevant fixes to it forever is our responsibility now that upstream has abandoned the linux-distro-bootstrapping use case. It is just unfortunate as that was supposed to be one of the goals.

genxy

14 hours ago

Frozen code seems even better for this use case.

nullsanity

a day ago

why can't you do:

C/C++ + libc ↓ LLVM/LLD ↓ mrustc + minicargo ↓ rustc 1.90 Cargo 1.90 ──→ OpenSSL ↓ Python x.py ↓ Rust 1.91+

lrvick

a day ago

That is what we do, and it takes several hours on most machines. Having to do that once to get mold and a second time to then build with the benefits of mold defeats any advantages for us.

n8henrie

a day ago

Agreed -- not my wheelhouse, but couldn't you just build rust earlier in your process using LLD, then build mold, then build everything else?

Would that meant that building rust is slower, but everything else is the same, and you don't have to maintain a fork of another complex project?

lrvick

a day ago

Rust requires python and perl and musl and openssl and 80% of the entire wall time of building the whole tree.

Rust is the single most expensive thing to bootstrap in any given linux distro.

It is the thing you need mold the most for to speed things up.

jiehong

a day ago

Hopefully, someone can do some work on decreasing the requirements needed to build rust itself and ease bootstrapping.

lrvick

a day ago

I sure hope so. Go is the gold standard to copy here, which is not surprising given the involvement of Ken Thompson who authored Reflections on Trusting Trust.

Unfortunately Rust does not even do full source bootstrapped deterministic builds themselves. Their release strategy is entirely the honor system, where everyone trusts their pinky swear that a single computer or person will never be compromised and allow for the injection of malware into the binaries everyone downloads from rustup.

The only reason we have even the crazy long bootstrap path we have today at all is the hard work of mutabah, an individual independent hacker that does it as a hobby.

The rust team seemingly considers memory safety as the only security problem that matters, and supply chain attacks out of scope.

as-182

a day ago

Because mold is faster. Linking is a significant bottleneck when building large projects.

mort96

a day ago

But we're just talking for bootstrapping Rust here.

someonebaggy

a day ago

The complaint is about bootstrapping, where speed is of little relevance.

lrvick

a day ago

It is of extreme relevance for us as we have to re-bootstrap every time we change any dependency of rust, and rust depends on almost everything that is expensive to build in a toolchain tree.

peterfirefly

a day ago

Use the mold binary from the last rebuild and only rebuild again if the new one is different?

jpollock

a day ago

If 80% of the tree is in the parent list for mold, the cache hit rate would be 20% - assuming the change distribution is even, which it won’t be.

MayeulC

a day ago

I have been thinking about Rust bootstrapping recently. Couldn't a Rust compiler without borrow checker be put together relatively easily?

Assuming the source code contains no issues (which can be checked later once the Rust compiler is built), one could leave that piece behind, and take the shortest path from .rs to executed code (C transpilation, or even an interpreter).

Orphis

a day ago

It seems like you would benefit from having reproducible and hermetic builds with a good caching layer so you don't do the same work again and again.

Then you can just use the latest built version to build mold and the next version of rust.

lrvick

a day ago

Our tree is entirely reproducible and hermetic. In fact we are the only distro that does this 100%.

The problem is changing any dependency of rust, even python or perl or musl or openssl or the llvm stack or any dependencies of the llvm stack have the potential of resulting in a different hash for rust.

Any time a dependency is changed all decedents must be rebuilt. Which we must do very frequently. That is where mold saved us a ton of time.

Orphis

a day ago

That's why you build a stage2 compiler at least with stage0 being the latest prebuilt.

Your stage1 will be based on whatever version is your stage0, but your stage2 should end up identical to any other build with any other stage0. So you don't have to rebuild the full chain all the time.

And then preferably you build a stage3 with FDO, you can easily get 20% more speed with a good corpus of tests.

Do you have at least a distributed cache with I guess sccache to save time for each rebuild?

lrvick

a day ago

If we added rust to our stage3, it would add 6 hours of wall time with the fastest 192 core CPU on the market to the time it takes people to reproduce and validate stage3. Add a couple days for someone with just a 16 core consumer CPU.

Our goal is to allow people to reproduce the entire tree from source in the shortest time possible, so no one has to trust us, thus encouraging as many people to do it as possible, thus preventing us from having the means to inject supply chain attacks without anyone noticing.

It is literally a goal for new users to be able to run a server that bootstraps itself, and then bootstraps and signs all future releases, with remote attestation proofs. The shorter that initial build window is the better, which is why things like fast linkers written in early-bootstrappable-languages are so important.

Orphis

a day ago

How long are all the rust builds in your distro? If the compiler is 20% faster (my LLVM Clang results), you "just" need 30 hours of building time to get even, and I'm sure there is enough rust around (maybe not all in your distro) to make this worthwhile.

In any case, people could be using the stage2 or stage3 binary and still get the same binaries in the end, they are just alternative versions of the same package that should be compatible.

And using a 192 core CPU is a bit futile, have you looked at how well the work can parallelize? Isn't there some good remote worker solution you could use to have more cores available for the whole fleet without locking a 192 cores machine for a single build? It doesn't look very efficient to me as it is.

lrvick

a day ago

Building from zero to rust is about 6 hours on my 192 core local workstation. The extra cores absolutely help. Each linux kernel takes 16 seconds. Rust still takes a few hours. And on a typical laptop, a day or more.

Building rust that far to build mold to then build rust again faster negates any benefits.

danudey

a day ago

Could you not either bootstrap rust with lld or bootstrap rust with mold2?

lrvick

a day ago

Bootstrapping rust with mold2 is going to be the likely play. My complaint is that now we must maintain mold2 forever now as mold3 rust edition is now incompatible with most of our tree which is made up of dependencies of rust.

biorach

a day ago

How much work is involved with mantaining a linker in deep maintenance mode?

lrvick

a day ago

Hopefully not much. I am just lamenting it is an extra dependency and chore we must carry and think about forever, and will have no official upstream support.

mohamedkoubaa

a day ago

Have you considered using Eurydice?

lrvick

a day ago

I am not finding any linkers by that name.

mohamedkoubaa

a day ago

It's a rust to C transpiler, I was thinking it could let you use a C version of the mold linker without maintaining one yourself.

lrvick

a day ago

Interesting idea! Will look into this.

how's bootstrapping rust story nowadays?

VorpalWay

a day ago

You could use https://github.com/thepowersgang/mrustc to get relatively modern Rust compiler (1.90 currently apparently, but every now and then that is updated). Then build newer rustc from there to reach the current 1.99.

But I don't get why some people are obsessed with bootstrapping. Yes it is good to be able to do it, but it isn't something you need to do regularly.

Especially since rust had a much better cross compilation story than C or C++ (not as good as go or zig though), so you don't need to bootstrap on a new architecture, just cross compile to it. Furthermore, new architectures for hosting a compiler (as opposed to just a target, like microcontrollers) is a rare event. Just something that happens every few years.

lrvick

a day ago

> But I don't get why some people are obsessed with bootstrapping.

It only matters if you have supply chain attacks in your threat model. Given they are up 400x since 2019, they should probably be in almost every threat model. Most distros operate on the honor system and that is not going to survive the post AI world.

Orphis

a day ago

That's why you have signed packages that everyone can rebuild and verify from the very early stages of bootstrapping to speed up the process and get back to a synchronization point that is reasonable.

Distros that are not fully hermetic and don't have reproducible packages (and there are many layers of reproducibility) will certainly have issues, but that's not a huge problem as long as you have a documented path to getting back to the current state. It doesn't need to be the fastest path, just a verifiable chain of trust.

lrvick

a day ago

We do sign our packages, but the threat model is to minimize the time it takes for users to not have to trust us.

It is critical to encourage many independent verifications that it be as fast as possible that someone can go from a clone of our tree of pure source code to the exact release hashes we publish.

Adding any binaries to that means someone that distrusts us must now go build those past releases as well, and if they rely on binaries, they must build those past past releases as well. This approach would make verification time go up dramatically every release.

Orphis

a day ago

You don't need to verify the chain of trust all the time, you can remember up to where you had trusted everything when you adopt your solution.

Also, while you can use the latest prebuilt as a stage0, you could also use any other prebuilt from earlier down the chain as a stage0 and still get the same stage2 binaries. This is the next trust checkpoint you have, and it should be fine to have multiple ways to get there too, one for quick iterative releases and one that is reusing the minimal amount of bootstrapped packages.

Or you just try to only upgrade your stage0 once in a while when extremely necessary and it would build your new stage2.

So many ways to optimize the system, it's a choice to refuse to reuse what was trusted yesterday in order to build the next stage.

ffaccount2

15 hours ago

>it's a choice to refuse to reuse what was trusted yesterday in order to build the next stage.

If the point is to let independent people verify the full build chain, then there is no "trust from yesterday" to reuse.

Even worse, if every release N depends on trusting release N-1, then you need to do a full rebuild of everything in the whole history, instead of just building the current release. This doesn't sound reasonable, or at least sounds incompatible with the project goals.

lrvick

14 hours ago

It gives me hope that at least some people understand the problem!

rui314

21 hours ago

That's unfortunate, and I'm sorry the Rust rewrite makes StageX's bootstrap more complicated. StageX has a fairly specialized requirement here, though, and I don't think it would make sense to keep mold in C++ just for that use case.

Using the final mold 2.x release as a bootstrap-only linker and then replacing it with mold 3 once Rust is available seems like the most practical compromise.

lrvick

6 hours ago

Yep that is what we are going to have to do. Is what it is. Thanks for maintaining an amazing linker, and we will surely ship v3 at the last mile for end users.

>Rust is not actually the right tool for all problems.

"my problems (that are not mold's) are not solved by mold. How dare mold make those decisions?"

Maybe you should rewrite more of your linux distribution in rust so it's available earlier in the build process and get back to it being the default linker.

jacquesm

a day ago

They were solved by mold. And then mold decided to un-solve them.

kelnos

a day ago

The mold developers are not responsible for what people want to do with their software.

No, they were never one of mold's goals. You don't get to assign solutions to people _and then blame them_. Hyrum's Law being a load bearing part of your infrastructure isn't Mold's problem.

jacquesm

a day ago

I can't really agree with that. If you ship your linker with a disclaimer saying 'don't use this for early stage stuff, one day we might decide to rewrite the whole codebase in a language that won't be available early enough' then you'd have a point. But if you didn't that is precisely the kind of use case where mold shines, and so inevitably it will be adopted. Your downstream should be precious. At a minimum you should communicate such intentions, do so timely and hear what others have to say, even if afterwards you decide to push on.

ndiddy

a day ago

Pretty much every piece of open source software ships with a legal document saying that the software is provided as-is without any warranty or guarantee of functionality (i.e. https://github.com/rui314/mold/blob/main/LICENSE ). Unless I had a support agreement with the maintainers that superseded that document, I would not expect any special treatment from upstream. Personally, whenever I add a new dependency, I do it with the knowledge that I may have to either replace it or take on maintenance myself in the future.

jacquesm

a day ago

Sure. But the ability to do this kind of rugpull is new and so you will get responses like these and that's perfectly logical regardless of the legalities. I'm not saying you can't. But I am saying that if you ignore your users - even if what you ship is free and open - you are still going to be subject to backlash if you change course very rapidly. You could also simply create a new project and support the old one during some transition period. Principle of least surprise and all that.

ndiddy

a day ago

I guess "AI-assisted rewrite in a language you can't use" is new, but it's functionally the same as "stopped working on the project" or "took the project commercial" or a host of other risks that have always existed with depending on someone else's work that they're providing for free with no formal relationship.

kelnos

a day ago

Are you seriously arguing that every open source project needs to anticipate every single way any of their users might use their software in the future, and also any future change they might want to make to the software, and write up a doc of disclaimers and caveats?

This is ridiculous. The amount of entitlement I'm reading here is gross. This is how you burn out maintainers and drive them away from open source.

> Your downstream should be precious.

No. Every open source maintainer is free to decide for themselves how much or how little they are willing to bend over backward for the sake of serving all possible user needs. If you don't like that, then feel free to build whatever you need yourself, from scratch.

jacquesm

a day ago

Strawman detected.

dwattttt

a day ago

Could you... go into detail about how?

> Your downstream should be precious.

Yeah, for stuff we wanna rely on, this should ideally be the default. I'd go one step further and say no breaking changes past 1.0.0 at all. Instead people should favor creating entirely new projects (forks or not) and jump over to those, leaving the old one behind, if they want to do massive changes to something.

Of course, no one would be forced to do this, but it feels like if more did this, long-term supporting stuff that depends on those things would be a lot easier, if things could just be instead of changing under our feet all the time. Thank god for Nix and NixOS, even with their warts.

3836293648

16 hours ago

It's literally a 3.0.0. Why have a major version at all in that case?

eviks

a day ago

What's that benefit of a new project and abandoning an old one when you can just as easily continue using the old version?

>Your downstream should be precious

Is my downstream paying me for the maintenance ?

If not, downstream may maintain mold2 themselves forever because their _extremely specific_ use case is neither a promise nor a valuable thing for mold to maintain. Debian understood this a while ago and isn't whining when they have to maintain their own fork. Maintainers maintain.

If enough downstreamers are unhappy about choices, they are also welcome to fork, until the base project is abandoned. Or they realize their usecases are extremely narrow.

(In addition, it's an extremely hypocritical and purist demand, because I am pretty certain that their distribution has, at some point, an arbitrary executable to make a compiler from. So they're probably okay with blobs, just not that one in particular.)

jstarks

a day ago

What? Why would mold be particularly valuable for early stage compilation, to the point that you’d explicitly cater to that user base?

someonebaggy

a day ago

I actually made a Linux distribution that relies on this being your latest comment, since you didn't say this wouldn't always be your latest comment. If you post any newer ones, it will break.

Edit: darn, it broke

ChickeNES

a day ago

Luckily Mr Clanker can do most of the work for you at least (I myself have already had Claude/Codex rewrite a couple of Rust projects in C, worked great)

senderista

a day ago

And repeat that work every time you sync with upstream?

DetroitThrow

2 days ago

I'm a bit confused why a decision to decrease the maintenance burden of Mold by switching languages makes Rust the wrong tool here? Fearless concurrency sounds like a huge benefit for what they're doing, given the resources they have.

compiler-guy

a day ago

It's simply an ordering problem. When building the entire world from scratch, usually the C and C++ toolchains are built near the very first, and Rust toolchains built somewhat later. Anything written in Rust must come after the Rust toolchain is built. You need a linker as part of your C++ toolchain, so it must be written in a language ready to go at that point. If it is written in C, you are done. If it is written in Rust you have to wait. So a Rust-based linker can no longer be used from that very, very early C bootstrap.

It's not a big deal for normal users, where you have Rust ready to go. Kind of a bummer in this case, but this is a specialized one.

veber-alex

a day ago

Can't you just use a statically linked mold binary for the initial bootstrap, just like you use some kind of pre-existing, basic C compiler?

compiler-guy

a day ago

You could. But now you are trusting an entirely new toolchain that can build the Rust-based linker, even if it is statically linked. This toolchain is much larger than a comparatively small and much more easily understood C bootstrap toolchain.

This effectively triples or quadruples (maybe even more) the amount of code you need to trust for the cold bootstrap.

Why do you keep saying C? Mold was C++ with some dependencies like TBB or zstd.

compiler-guy

a day ago

Too many years at Google where the two terms are most often used interchangeably; which is imprecise, but most of the time it doesn't matter. Forgive me for being imprecise here.

The fact remains that rust adds yet another toolchain to the boot process that needs trust and verification.

someonebaggy

a day ago

Maybe that's about to flip. Maybe Rust will be built first, followed by C.

lrvick

a day ago

Sorry to break it to everyone but Rust depends on python which depends on perl and openssl and the majority of any lean full source bootstrapped toolchain tree.

Unless the gcc rust engine is mature any time soon (lol), we have no path to use rust until very late game in a distro build.

The earliest we can bootstrap a go compiler is about 10 minutes. It builds directly from tinycc. Add 6 hours for our fastest compile of the shortest path to rust, with 192 cores.

I am a rust fan too, but it is the worst language to bootstrap, which is why for systems programming I still must often revert to C to have a small and reviewable and fast to build dependency surface.

afdbcreid

a day ago

Rust requires LLVM (or GCC) which requires C++, not just C.

VorpalWay

a day ago

LLVM is the default yes, there is also an experimental cranelift backend, which is all Rust. Not sure if it is good enough to build the compiler itself using it.

(There is also a GCC backend called codegen_gcc that is pretty far along. And a separate reimplementation of both the frontend and backend using gcc and C++, called gccrs, which is not nearly as far along.)

afdbcreid

a day ago

I'm pretty sure it does not support enough features to build the compiler (e.g. inline asm is not supported at all), and it will also be very, very slow to unusable.

ahartmetz

a day ago

Cranelift is supposed to compile without optimizations (-O0) faster than the regular Rust compiler, and supposedly, the code runs at about the same speed as regular -O0 compiled code. That generally means 2-5x slower runtime performance, which is just fine for bootstrapping.

afdbcreid

12 hours ago

I've never tried to bootstrap rustc with cranelift. I think some people did, not sure what were the results. I do know that bootstrapping rustc with -O0 is so slow it's essentially impossible even for just one bootstrapping step. -O1 might be possible.

someonebaggy

a day ago

Maybe it should be rewritten in Rust then

tadfisher

a day ago

One of the explicit goals for Mold 3.0 is to promote its use as the default linker in Linux distributions, and bootstrapping complexity is certainly worthy of consideration in that realm.

> We will then conduct extensive compatibility testing and work closely with Linux distribution developers to make it practical for them to adopt mold as /usr/bin/ld. Making this happen is one of our highest priorities for mold 3.x.

lrvick

a day ago

We are one of only two distros that uses mold as the default system linker. Unfortunately that will be stuck at mold2 until the full tree is built, but we can swap to mold3 for end user consumers of the tree. Sucks that we have to maintain and use both now.

duped

a day ago

Why do you need to bootstrap a toolchain to bootstrap a distro?

I know this is common but it seems like either an aesthetic decision, or glibc cruft.

imoverclocked

a day ago

There are classes of virus that are hard to detect. One is a compiler virus that passes itself from compiler to compiler. You only get rid of the vector by bootstrapping from 0.

Aissen

a day ago

No, you can do bootstrapping and save binaries for reuse with hash verification. Android did that for its Rust toolchain: https://cs.android.com/android/platform/superproject/main/+/...

Bootstrapping at every build does not save you from the threat you think it does.

lrvick

a day ago

We only re-bootstrap the layers which had dependencies change under them. Early parts of the tree rarely change so we often do not have to rebuild these across releases, but very late tree things like rust depend on almost everything and something in the rust dependency graph changes almost every release.

Using binaries from past releases is a strict downgrade in terms of verification speed, as it means a new independent reproducible build verifier must now build both trees, doubling the release verification time, and erasing any wins mold3 could otherwise offer.

Google can rely on lots of centralized internal provenance tooling to prove cached binaries are not tampered with to other Googlers but when the goal is proving end to end full source bootstrapped build integrity to any interested user from the public in the least time possible, the requirements are significantly higher.

dwattttt

a day ago

> prove cached binaries are not tampered

Is this not just verifying hashes? What further effort do they go to to prove a binary hasn't been modified?

duped

a day ago

Sure but that's a compiler bootstrapping problem. It doesn't answer the question: why do you need to bootstrap the toolchain to build the distro? You can reuse a trusted toolchain that's been safely bootstrapped .

lrvick

a day ago

Because no other trusted toolchains exist under a threat model that trusts no single person or computer. We -are- the trusted toolchain.

duped

15 hours ago

If you cant trust your own machines outputs to build the next version of your code why should any user trust them

lrvick

14 hours ago

No user should trust them. If someone threatened me to put in a backdoor, I would. But it would not matter. Trust happens when many independent signed builds from reputable people all get the same result. That is why we make it easy and inexpensive for users to build the whole world from 0 from source at any given commit.

nicce

a day ago

If the need is well justified, maybe there is great chance to ask adding #[no_mangle] and extern C support? Since release is very fresh. If that is causing the problem. Or is some dependency the issue?

afdbcreid

a day ago

The need here is bootstrapping. If mold is written in Rust you cannot even compile it.

nicce

a day ago

Even Rust is written with Rust. But yeah, if you want to bootstrap absolutely from the beginning, I see the issue.

lrvick

a day ago

We regularly and redundantly full source bootstrap our entire tree from 186 bytes of machine code due to a supply chain security policy that forbids trusting any single human or computer.

thesz

a day ago

  > 186 bytes of machine code
This is better (smaller) than Forth!

Aissen

2 days ago

I expected this to be a multi-months rewrite, not 3 weeks. I almost forgot we live in the agents era now.

Edit: this seems to have been cooking for a while when the first commit dropped: https://github.com/rui314/mold/commit/f41bfcd5c72ca30cce6498...

rui314

21 hours ago

When I announced the rewrite, it was actually already mostly done.

AI's coding ability is truly amazing. We've spent decades inventing languages, tools, and methodologies to help us write better code more efficiently, but AI-assisted coding is the biggest breakthrough in programming productivity I've seen in my lifetime. It's genuinely incredible.

Was the conversion assisted? I get the sense that Mold's author is an extremely competent guy and it would be quite feasible for someone like him to do it on his own.

SuperV1234

a day ago

He decided to use LLM assistance exactly because he's an extremely competent guy.

ahartmetz

a day ago

It's also his third(!) time writing a linker: lld, mold 1-2, now mold 3. mold 3 is also supposed to do what he never did before: implement all of the GNU ld features (notably linker scripts, I think it also has some optimization features not available elsewhere) so that it can finally be replaced, never needing it as a fallback anymore.

afdbcreid

a day ago

Assisted, yes. Claude is a coauthor for some of the commits and I think he also said that explicitly. Vibe-coded, he claims not.

awoimbee

2 days ago

So what differentiates mold from wild now ? Is wild using different data structures?

(Wild is another fast linker, that only supports Linux)

vlovich123

2 days ago

Wild's primary purpose is incremental linking - I guess now the primary difference is a race + subtle differences in performance between projects.

weinzierl

a day ago

wild's primary purpose is incremental linking, which it incidentally does not support at all, while still being faster than mold by a long shot.

dralley

a day ago

They're not that far apart anymore. Wild is still faster on average but they trade blows occasionally. Both are well ahead of the next best.

avadodin

5 hours ago

Maybe the lead can use the time he didn't use to translate mold into rust to write chibirustc for fast local rust compilation and we can all forget this was ever an issue.

sharktheone

a day ago

Oh what? That was not on my bingo card for 2026. Mold already was incredible when it still was in C++ and I assume it is even better now that it is in rust

Pleasantly surprised to see a ~600 lines Cargo.lock (vs 2200 for wild, for example) here! Hope it can slim it down further.

Something I really look into because the crates situation (due to a too small stdlib à la R5RS Scheme) is Rust's Achilles' heel, since it directly affects its security claims.

WhereIsTheTruth

2 days ago

It was transpiled, not rewritten, rewrite implies by hand

tialaramex

a day ago

If it was transpiled there would be what GNU calls a "preferred form" of the linker source in which it wasn't written in Rust.

This is true for - as an example - the WUFFS GIF decoder. You can get C which decodes GIFs and was transpiled from WUFFS, but that's awful code and nobody wants to modify that code, whereas the WUFFS source code for the decoder is fine.

When we look at rui314's changes to mold today after 3.0 release, they just modify the mold source code in Rust, as you'd expect if this was in fact now written in Rust.

Dylan16807

a day ago

> If it was transpiled there would be what GNU calls a "preferred form" of the linker source in which it wasn't written in Rust.

No, that's only sometimes true. If you like the transpiled version and start editing it directly then that's the new preferred form, but it was still a transpile to get from one language to the other.

dcre

2 days ago

No it doesn't.

mi_lk

2 days ago

Damn. I thought Zig would be a perfect rewrite language for Mold since it's a better C in many ways

vlovich123

2 days ago

Zig makes you choose either safe or fast, not both. With Rust you can get both (generally).

senderista

a day ago

Part of the safety you get with Rust is runtime bounds checking, which is equivalent to Zig in "safe" mode.

vlovich123

a day ago

Bounds checking is like 1% of the safety that Rust guarantees you. That's a false-equivalence. The bigger ones are around memory safety like use-after-free, double-free, and memory issues in the face of concurrency. Zig does not help you with that.

O3marchnative

a day ago

There are many ways to get rid of bounds checking in Rust. One of those ways is just utilizing iterators. For alternatives, there's a bounds check cookbook you should checkout [0].

Thus far, I've written two libraries that utilize Rust, and they are usually associated with "high performance". I was able to exceed the performance of similar libraries written in C/C++.

To be perfectly clear, this doesn't mean Rust can replace C/C++ in every scenario, but the actual perf gap is narrower than some folks may be willing to admit.

[0] https://shnatsel.medium.com/how-to-avoid-bounds-checks-in-ru...

afdbcreid

a day ago

Bound checks, which are part of what you get with Rust, also exist in ReleaseSafe Zig, that is true. But if that is what you wanted to say, your statement is very confusing and also does not answer the GP. If you meant to say ReleaseSafe and Rust are equivalent, then this is just false. ReleaseSafe still has many UBs (most importantly use after free and double free), not to mention that Rust solves many problems at compile time and Zig only at runtime.

senderista

a day ago

Sorry, I agree I was unclear and your restatement accurately reflects my intent. My biggest concern with Zig safety is that ReleaseFast turns asserts into assumptions, which is equivalent to injecting UB at every failed assert site. I doubt many users are aware of that "feature".

Zambyte

a day ago

Zig makes you choose that at build time. You can use safe builds for development to catch bugs, and fast builds for release. You can even use safe builds in of the release, and optimize hot paths with fast builds.

In practice, projects written in Zig very much can choose both.

bhaak

a day ago

You make it sound as if all bugs could be catched in development builds.

Zambyte

a day ago

Saying that sounds like you're claiming Rust can catch all bugs at build time. Obviously neither can.

pjmlp

a day ago

Still doesn't have a good answer for use after free, although the allocators on the last release might help into that regard.

audunw

a day ago

That’s not entirely true. There’s a third choice. You can go the TigerBeetle route (TigerStyle). Zig is extremely well suited for software where you just avoid dynamic memory allocation all together. It’s extremely fast, very effective and arguably very safe. But not suitable for all applications obviously.

There’s also another caveat that there’s an effort to build in build time static memory safety checks in a way that’s more general than Rusts borrow checker, through a kind of plug in system rather than forced into the language. It’s not part of mainline Zig yet but this seems to be the direction Andrew wants to go.

https://github.com/ityonemo/clr

p-e-w

a day ago

When choosing a language, people don’t only look at language features but also at community adoption, the library ecosystem, industry backing etc.

Zig isn’t even in the same league as Rust regarding these things. Zig may still be around and active 10 years from now, Rust is guaranteed to be.

bryanlarsen

a day ago

Additionally, Zig just released 0.17, Rust's 1.0 was 11 years ago. The choice may have been mostly for non-technical reasons.

Ar-Curunir

a day ago

Indeed part of the choice is because Rust is now an established systems language.

uncle_kostya

2 days ago

I'm curious about motivation - bounds checks for corrupted inputs seems like it would be one, but it also seems that fixing corrupted input handling in a C/C++ code base would not be too hard, and probably less of an effort? So why did you choose the rewrite?

And second, did you use any AI tools for the rewrite?

fotcorn

2 days ago

> fixing corrupted input handling in a C/C++ code base would not be too hard

The best programmers on the planet have tried and failed with this task for 50 years now, so I don't think this is true.

The main disadvantage of Rust right now is not supporting some more obscure platforms, but because mold wouldn't support them anyway I don't see that as a problem.

embedding-shape

2 days ago

> The main disadvantage of Rust right now is not supporting some more obscure platforms

Last time I checked, I got impressed by the wide platform support, once you go down the tier list (https://doc.rust-lang.org/nightly/rustc/platform-support.htm...). What "obscure platform" specifically are you thinking about, that is currently missing from those lists?

fotcorn

2 days ago

Just to be clear, I think this is a very small disadvantage. GCC and therefore C/C++ supports some old stuff like SuperH, Intel Itanium, PA-RISC and a bunch of microcontroller archs that LLVM does not.

However, there is now a Rust codegen plugin for GCC, so even this disadvantage is now basically moot.

Rewrite all the things!

afdbcreid

a day ago

The GCC backend is fairly complete but still does not support everything.

jcranmer

a day ago

The list of supported GCC architectures is here: https://gcc.gnu.org/backends.html

LLVM doesn't have an equivalent annotated list, but it is missing alpha, bfin, c6x, fr30, frv, gcn, h8300, ia64 (aka Itanium), lm32, m32c, m32r, mcore, mep, microblaze, mmix, mn10300, moxie, nds32, nios2, pa (aka PA-RISC), pdp11, pru, rl78, rs6000, rx, sh (aka SuperH), storm16, v850, vax, and visium. For its part, LLVM does have some targets that GCC doesn't have (mostly related to GPU compilation).

The "big" targets that GCC has that LLVM lacks are Alpha, Itanium, PA-RISC, and SuperH, with Itanium being sufficiently weird that it's pretty firmly in the "fuck this" category from a maintainer's perspective, and people are actively ripping out support for it.

torginus

2 days ago

yeah this makes no sense to me. Why would Rust limit the LLVM backend's ability to generate code for a particular platform?

panzi

2 days ago

I think the claim is that GCC supports a few obscure platforms that LLVM doesn't. Don't ask me which, that is just the claim that I heard multiple times. So it's not Rust that limits platforms, but LLVM. And they say the GCC backend efforts of Rust are meant to deal with that.

bkallus

2 days ago

> Don't ask me which

There are many. Two big ones are Alpha and PA-RISC. NetBSD and Linux continue to support both. Linux distro choices are pretty much limited to Gentoo though.

steveklabnik

2 days ago

1. You've got the causality backwards. LLVM backends don't come for free, you have to write code to enable support for them.

2. LLVM does not support as many backends as GCC does, so even if you did get 100% of the LLVM supported backends up and running, you'd still be missing some.

VorpalWay

a day ago

Worth adding to this: there may also be code needed in the rust part (not just the llvm part) to enable support for a target triplet. There certainly will be for the OS and libc part of the triplet: what specific syscalls exist and should be used to implement the standard library on a particular flavour of BSD etc.

But even for the architecture part there can be differences. For example, my understanding is that a lot of the calling convention details need to be handled by the frontend in LLVM, leading rustc to duplicate logic from clang here. Things like varargs FFI with C code can be particularly gnarly.

saghm

2 days ago

My very naive understanding is that a part of what makes mold fast is concurrency, which I'd expect to be a lot more error-prone in C/C++. Not having to worry about data races might give more confidence with trying out more complex techniques for how to split up work in a way that ends up making things faster

pornel

2 days ago

I suspect fearless concurrency is another motivating factor. Better perf can be achieved by squeezing more parallelism, but without borrow checking it's difficult to do fine-grained parallelism correctly.

wpollock

a day ago

Is "fearless concurrency" a technical term? I thought it was just the catchy name of a chapter of the Rust book.

LoganDark

2 days ago

> And second, did you use any AI tools for the rewrite?

IMO, given the recent commits: almost certainly.

someonebaggy

2 days ago

It's possible to write safe code in C or C++ but it's extremely difficult to read existing code and prove it's safe, without using as much effort as it takes to write it in the first place. This includes the code you wrote last month whose surrounding code has changed. And you have to be right every time while the attacker only needs you to be wrong once. The problem is not writing the code, it is continually verifying it.

VorpalWay

a day ago

As someone who worked professionally with both C++ and Rust, I mostly agree with this, but I would say that writing the concurrent code in C++ is also hard.

The only way I found that works reliably is stick to a small well defined set of mostly safe primitives. E.g. at work we use message passing / event bus architecture everywhere, which works great for a robotics / industrial context. But even then, if you somehow mess up and have a variable accessed from event handlers in two different threads, it is tough to spot other than if you get lucky and observe it with a build using TSAN.

With rust that class of mistakes is just entirely eliminated, which makes it easier to to concentrate on the hard things that actually matter (like the domain specific logic).

nine_k

a day ago

> in a C/C++ code base

It's like "in a Zodiac boat / aircraft carrier navy". This customary putting C and C++ into the same bucket is as amusing as it is unproductive.

flohofwoe

a day ago

Apparently the old mold version had a couple of plain C source files in the source tree (both vendored 3rd party libs but also in the 'regular' source code), so technically C/C++ is correct in this case ...and looks like the Rust version also has some (very minimal) C code left (looks like mostly varargs stuff, I guess Rust doesn't have a concept of varargs?), so it could be called a C/Rust project ;)

https://github.com/rui314/mold/tree/main/c

pjmlp

a day ago

Unfortunately too many folks still do C style programming in C++.