Back to blog
AI governancedialectslanguage modelsEU AI Actlinguistic diversitylow-resource languagesfairness

Where to draw the line: Dialect representation

AI Usage
13%
GlennSeptember 7, 20267 min read

Where to draw the line: Dialect representation

Åse Wetås, now head of the national library after years of running Språkrådet (Norway's language council), said something along the lines of "we want to build a model that represents all Norwegian dialects" (this is paraphrasing) in a talk she gave about AI as Infrastructure at CuttingEdgeAI: Sovereign AI Infrastructure earlier this year. But with Norway having hundreds upon hundreds of regional dialects (a recent claim says over 1300), is this even possible? Do we include the main ones, or should we try to represent every single regional difference from town to town? Exactly where is that line drawn?

This question came up again when Christopher Bosley presented "Ethnonationalism by Algorithm" by Spencer Overton at the AI Governance Club a few weeks ago. Overton points out that governing is not unbiased, and that we need a framework that will combat this. One part of his argument is what he calls "homogenization": LLMs as averaging machines, optimizing for dominant linguistic patterns and marginalizing everything else. That left me with a thought that has been in the back of my mind for a while. When is dialect representation good enough?

Back to Åse's talk. At the end I raised my hand and challenged her on the idea, asking what their strategy was. Because I would argue that representing all dialects in a meaningful way is impossible with probabilistic technology. Low resource languages are already struggling, thus small regional dialects do not stand a chance. At least not yet. This is especially true in Norway, where there is a distinction between the "written language" and the "everyday spoken language". And even though we have two official written languages, there are thousands upon thousands of nuances in the spoken dialects that are not properly represented. It is not only because of the lack of representation, but also because of the curated data that is going into training the models. We want good quality data, and that sometimes means dropping comments under news articles, forum posts and general chatter, and prioritizing well-structured documents. This in itself will miss the mark, as by taking out how people casually write (and speak) to focus on "well-structured, curated data", we delete some of the diversity by design.

She didn't push back. Wetås isn't a technical person, she's the head of the national library with a vision for what this should become, and she said as much: she didn't know enough about the technical side to give me a good answer, but her ambition stayed the same regardless, something I find honorable (though perhaps a tad unrealistic).

NorMistral, one of the Norwegian language models developed outside the National Library, documents this constraint openly: Norwegian and Northern Sámi at the core, with Swedish, Danish, Icelandic, and Faroese folded in for what the researchers call "knowledge transfer," because there simply is not enough Norwegian to work with on its own. The National Library's own Borealis family takes a different approach, built on Google's Gemma 3 and fine-tuned on Norwegian instruction data, with access to a small amount of copyright-protected Norwegian press material through an agreement between rights-holder organizations and the Norwegian government. Two different strategies, same underlying constraint. So here we are, building models for every Norwegian dialect, while the base architecture was trained in California.

Should we still try?

I find the pursuit impossible with today's technology. That does not mean I think we should stop working on it. Good science includes pushing boundaries and looking at problems from different angles.

What I do not think we should do is put all our eggs in that basket right now. A small number of people should keep working on the technical problem. A more pressing matter, in my opinion, is to figure out where to draw the lines in the sand to govern what already exists, whether or not it ever gets solved.

So what IS good enough?

Nobody has answered this question cleanly. And the more you look at it, the clearer it becomes why.

The problem starts with definition. Language and dialect exist on a continuum, not in boxes. Norwegian and Swedish are technically separate languages, but many Norwegians will understand a Swede from Gothenburg better than a Norwegian from some remote valley speaking their local dialect. The boundary between language and dialect is political as much as it is linguistic. Which means every governance framework we have is built on a category that does not hold up under scrutiny.

The EU AI Act is the most detailed piece of AI regulation that exists. It spells out exactly which demographic characteristics a dataset needs to be representative of, things like health, safety, and non-discrimination between women and men. It does not mention language once. The most exhaustive AI law on the books hasn't even reached the question I'm asking about dialects. And yes, you can technically chuck something into training data and call it representation. But when a dialect produces almost no written text, the training signal for that dialect is sparse, and what little exists gets averaged into a vastly larger dominant distribution. If complying with a legal requirement simply means adding data that will be lost in latent space, it is not really worth governing at all. The same issues remain at the language level. How are we supposed to weigh every language equally when meaningful representation is a huge technical challenge that we have not yet solved?

This brings us to the question of thresholds. Is there a number we can agree on?

Attempted answers

In 1978, the EEOC settled on its four-fifths rule for hiring and promotion decisions: 80%. So should we use that as a basis and agree that if 80% of the dialects are 80% represented 80% of the time, we can call it a success? The problem is that applying it to dialect representation requires you to answer three questions before you can even start. 80% of what? Measured how? Verified by whom? Remember that this is a floor to be raised, not a ceiling to crash into.

A paper called "Fair Enough?" (Regoli, Castelnovo, Inverardi, Nanino, Penco) makes almost the same argument about fairness generally that I'm making here about dialects specifically: that "fair algorithms" is close to a meaningless requirement on its own. It only becomes actionable once a society actually decides what fairness means in that specific context, and right now, nobody has. The paper doesn't propose a number either. It just maps out every question we haven't answered yet. You cannot solve a measurement problem without first solving the definition problem.

So where do we draw the line on fairness? Is there a fairness threshold we can push against, or is it subjective for each model provider? And if it is subjective, does that mean it is ungovernable? That might be the uncomfortable truth.

When it comes to dialects, I really think we should aspire to include as many diverse dialects as we can, but understand that the dominant written variety will generally remain best represented in the model. Prompting may elicit a smaller dialect more effectively, but it cannot manufacture linguistic knowledge that the model never acquired. In that case, the smaller the dialect, the less reliable that adaptation becomes. And even if we are to aspire to represent all dialects, we cannot put it into law yet, because it will be impossible to comply.

Until we can reliably teach models to preserve extremely rare linguistic patterns without the dominant distribution eating them, there are no two ways about it. We will not be able to reach Åse's idealistic vision. We will have to treat "good enough" as an open question, and, for now, accept that the statistical distribution of language will not give us perfect diversity.

Where to draw the line: a blog series

This is a new blog series I'm writing that will explore when and what we should actually govern when it comes to AI. I will be writing about digital likeness (a question I already touched on in my Digital Doppelgangers piece), data hoarding, autonomy, acceptable risk, and more.

Further reading

Dialects and representation in AI

Low-resource languages in AI

Norwegian and Nordic language models

Fairness thresholds and governance

Scandinavian mutual intelligibility