hypfer
6 days ago
I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms.
The kind where malicious intent is okay if the words are nice.
___
Or, rephrased: How big is the space in which you can tune this model without retraining.
Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed?
__
Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?"
That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off.
___
Edit: Looking at the paper though.. probably not.
I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
xp84
5 days ago
I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the thinly-coded language, so why are we censoring everything in the first place?
Gigachad
5 days ago
The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened.
Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.
wongarsu
5 days ago
Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate
gwerbin
5 days ago
That's the point
gwerbin
5 days ago
Moderation by well-intentioned humans is totally different from moderation by a large corporation with a profit motive and political pressures.
torginus
5 days ago
Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.
hypfer
5 days ago
Ha, thanks, I completely forgot that that was a thing. I got hit by that too, which was incredibly funny and a great throwback to early 4chan culture.
So fwiw, these things at least do breed creativity. Same as with the aforementioned "unalive" or "keep yourself safe".
There's some beauty in the online hellscape if you just go looking for it.
torginus
5 days ago
Yeah another kind of a*hole who I've encountered on Reddit were the ones who would keep mouthing off to you (carefully keeping within 'letter of the rules' of moderation), trying to goad you into snapping at them, and they'd insta-report and ban you.
Though for this scheme to work, it required Reddit mods to be... Reddit mods(can't come up with a better insult), at which point the whole thing seems so pointless - why have this elaborate song and dance with rules you pretend to follow, when you can and will ban anyone who rubs you the wrong way. Just announce that the rules are whatever the mods feel like that day, and stop pretending.
user
5 days ago
inigyou
5 days ago
When did that feature get taken offline?
hnbad
5 days ago
The problem is that you can't have content moderation without having editorial intent. There is no such thing as "neutral moderation".
Most moderated spaces these days rely on moderating based on "civility" because it can be excused with jargon like "creating a marketplace of ideas" which ignores the reality that the scope of discussions selects for who participates in them as much as the way they are phrased - a zebra will be less inclined to participate in a "marketplace of ideas" where a recurring topic of discussion is how zebra meat is best prepared for consumption even though that space might be very attractive to lions and tigers.
But I'm not sure if this is truly accidental. Ever since the advent of online advertising, online spaces have been overtaken by corporate interests. Heck, it's even endemic to "social media" given that those platforms themselves have turned into major corporations or at least were acquired by them. I'm not implying any nefarious intent but "civility" is certainly the dominating factor when it comes to what corporations care about when it comes to content moderation - anything beyond that is largely about what target demographic they're trying to attract and what virtue/vice signalling is optimal based on the current social and political environment (cf. various major corporations demonstratively dropping "DEI" initiatives following Trump's election).
I would also argue that in terms of content (rather than tone), corporations are necessarily also much less tolerant of "far left" issues than "far right": all social justice movements at the end steer towards anti-capitalism because they run counter to the perpetuation of social (or economical) hierarchies. This is why we saw so many tech companies (including those formerly described as "very liberal") shut down their DEI initiatives even before Trump got elected - because the political window had shifted to the point where this had become defensible while at the same time many DEI ideas had become so widespread culturally that these initiatives now became a direct threat to the "(old) white men" running those companies. This had been inevitable but DEI was seen as a necessary marketing effort (both internal and external) at the time, not something truly adopted on ideological grounds. This is also why speakers/trainers promoting "white guilt" were more popular - making your white employees feel bad is less threatening than making your marginalized employees think critically about the structures that lead to their marginalization; you want to individualize the problem, not direct attention to the systems underpinning it.
Another factor is that on social media content is mostly moderated "softly" by the algorithm. This is the content moderation you don't get to see because it can exercise editorial control where the simple word filters can't. The word filters create plausible deniability: they "try" to filter unpalatable subjects but those darn kids are just so clever and circumvent it. Meanwhile the algorithms can be fine-tuned so the topics you really don't want to see discussed stay off most people's "for you" pages - or even so those who would be attracted to them still see them and feel elevated and heard despite actually being isolated into their own echo chamber.
rancar2
6 days ago
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custom ML engine to handle our levels of traffic for reviews, which was among the largest in the US typical ranking top 3 on Google for the domain keywords. Back when BERT was the edge, a policy-adaptive model like this one from Mistral would have been an incredible cold-start solution. Most sites never have the massive volume nor budget needed nor skillset needed before you can train domain-specific models that outperform OOTB solutions. Generally, most people and site mean well and try to do well, so empowering those people with models like this can help the collective in my opinion, so I’m happy to see this released in this manner myself.
kerisi
6 days ago
[flagged]
rancar2
6 days ago
A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather benign French translation of “Enseigner la discrimination" which I also see awkwardly translated in the paper as well.
joshhart
6 days ago
Discriminative has a meaning in machine learning that I think is relevant here. There are "generative" models like LLMs that are learning joint probabilities P(X, Y) and "discriminative" models like logistic regression that learn conditional probabilities P(Y | X)
NopIdoN
6 days ago
"Discrimination" is exactly correct. What you suggest changes meaning.
rancar2
5 days ago
Thanks for the clarification. I’m very happily wrong here, and I appreciate the correction. This does emphasize from an editorial review that the double meaning can be distracting for someone who isn’t very deep in the terminology of this particular area, so it may still be worth rewording the subheading for a broader audience which will be interested in this model.
hamper653
6 days ago
> In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment.
In French as well. But it’s obviously not what’s meant here.
nikcub
6 days ago
> which seems to be mistrals whole thing
They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche.
Before the datacenter deals their revenue was higher than xAI's
There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost.
Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.
bunyak
6 days ago
[flagged]
nikcub
6 days ago
public ones are content moderation as above and previously llama guard, et al
OCR is also another field - Mistral have a model, so do deepseek
The ones I have experience with where you fine-tune smaller / faster models for business tasks like content writing, support, etc. by their nature stay private
Fnoord
6 days ago
Voxtral [1], maybe Robostral? I just saw on their blog they released OCR 4. Their is usually informative [2]
nolok
5 days ago
Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory).
So the goal is probably to be able to tune this basis to your ruleset.
As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obvious differences: hate speech and sex). Mistral is an european company with a french basis, so I doubt they didn't plan for that (otherwise they're complete morons, which I don't think they are).
egorfine
5 days ago
It sure feels to me an LLM company stands no chance in the industry unless they succumb to specific North American moral values.
AdamN
5 days ago
There are two usecases: 1/ general purpose, 2/ customer controlled.
Mistral is focused on the second one and every customer, whether it's Boeing or Airbus or Nokia or Glencore, will want lots of control over 'their' models. That's not a North American moral values thing.
For the first one yes there will be at most 2 or 3 model cultures but even there I think some customers will want really open and some will want more walled gardens.
nozzlegear
6 days ago
> Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Isn't Mistral a French company? Not that the French can't do cultural imperialism either, but they (the French) don't strike me as very SV.
asveikau
6 days ago
> "that one moderation style" we already know from current big tech platforms.
> The kind where malicious intent is okay if the words are nice.
Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
inigyou
5 days ago
It's the one where you get censored for using words like "kill" or "shit", but if you say "I'm going to come to your house and unalive you" nothing happens because you didn't trip the word filter. When the company gets in trouble for this, to fix it, the company starts removing all comments that include the word "house".
numpad0
5 days ago
I think you're referring to a different part of the same phenomenon as GP - the current big tech moderation feels a bit like using the shape of a crescent moon shape to cover a square.
dannyw
6 days ago
I thought OP was referring to AI platforms.
For example, by positioning what you’re doing as an accessibility tool and using the right words, you can get the latest models to write incredibly powerful malware without safeguards kicking in.
Fnoord
6 days ago
if you pull stuff like that off, and it gets flagged, it is obvious you are trying to game the system. Same with a system like this. It could be used in addition to human moderators, where the human moderator have to meta-moderate the AI's work. This could also be used as training exercise for new moderators. Then, the good and experienced moderators have more time to spend on edge cases, complex cases, fine-tune the AI, etc. In other words, it is able to do the most boring things to you.
And EU regulate, I mean as a counter example: we are not in panic about a nipple. When I was in a large museum in Paris, multiple women were breast feeding their infant. And why not? Kid's gotta eat. I'll refrain from insulting any world leaders, too easy, but you know many examples are available there regarding censorship.
Finally, it can take that BS argument away of 'oh we don't have manpower to moderate'. That is a low blow, too, by large commercial entities who could, you know hire and train? However, even a small company with not much money to burn could -in theory- win here.
I'd give this model a chance, if not only cause I've been impressed by Mistral past years. Yes, Le Chat / Vibe probably lags behind, but something like Voxtral (real-time and transcribe) is neat, and efficient.
charcircuit
6 days ago
It sounds like it is. You have a set of moderation policies and then you evaluate the model 1 time per policy if it is violating it. Then you combine the results into a score you use for taking actions off of.
HawtAds
6 days ago
> Kinda like cultural imperialism but with an ethical spin.
As the old adage goes, US innovates, China imitates, EU regulates.