Aksarium Codex Lexpressia Where the data stops
A walkthrough in three parts

Where the data stops

Platforms moderate by language. This is a walk along one language until the moderation runs out, and then an argument about why it ran out there and not somewhere else. It takes about fifteen minutes.

Arrow keys or space to move; M for this menu; 1, 2 and 3 jump to a module. Companion reading: the Bengali resource page.

Module 01

The line and the ledger

Myth we will unpick: a language is a thing you can hire moderators for.

Bengali has around 285 million speakers and sits sixth in the world. It is also a low-resource language, and it is not one language in any sense a classifier could use. This module walks it from end to end.

Where the maps put the border

In the Linguistic Survey of India of 1903, George Grierson grouped the dialects of Chittagong under Southeastern Bengali, alongside those of Noakhali and Akyab. Akyab is the colonial name for Sittwe, the capital of Rakhine State in what is now Myanmar.

A century later, the international standards register still agreed. Rohingya and Chittagonian shared one ISO code until 2007, when cit was retired and split into ctg and rhg.

The separation of Rohingya from Chittagonian is nineteen years old. It is an administrative act, not an ancient finding; before it, the register said what Grierson said.

Walk the line

Five stops, north-west to south-east, from the Hooghly to the Naf. Choose one and see what has been built for it.

Somebody did test the assumption

A single classifier over that whole line assumes the line is close enough to itself to count as one language. Meta has never published a test of that. The humanitarian sector ran one, because lives turned on the answer.

In 2018 Translators without Borders surveyed more than four hundred Rohingya adults in the Cox's Bazar camps, where the response had been built on the premise that Chittagonian was near enough to Rohingya to be understood. More than a third could not understand a simple sentence in it.

Two-thirds of the camp population have no formal education and about 66 per cent cannot read or write in any language. The preferred written language in the camps is Burmese, because the state that persecuted them is the only one that ever taught them to read.

The word is the argument

Myanmar's state discourse calls the Rohingya Bengali, and it does so in order to conclude that they came from Bengal and are therefore not indigenous. The 1982 citizenship law makes indigeneity turn on presence before 1823, so the whole question of who these people are is conducted as an argument about a date.

Calling the language Bengali is not a classification error. It is the argument, and the reference works have been supplying it with cover since 1903. When the Arakan Army displaced the junta from northern Rakhine, it kept the ban on the word Rohingya. The naming policy turned out to be the more durable institution.

That is the line. Now the other end of the same failure, where the language was not the problem at all.

Module 02

What the machine could not read

Myth we will unpick: the failure in Myanmar was that the model was bad at Burmese.

It was worse than that, and more specific. For most of the relevant period the platform could not reliably compare two strings of Burmese text, and the people best placed to report incitement could not read the form for reporting it.

The country that never standardised

Facebook's own engineering team, writing in September 2019, called Myanmar "the only country in the world with a significant online presence that hasn't standardized on Unicode". The dominant encoding was Zawgyi, which uses multiple code points where Unicode uses one, and permits variable vowel ordering.

The consequence, again in their words: the same word can be encoded as more than one sequence, so string comparison fails within a single document. Their own analogy is that CAT and CTA read the same. And this, they wrote, "made it hard to train our classifiers and AI systems to effectively detect policy-violating content".

The search that misses

Six posts. All six contain the same Burmese word, and all six render identically on screen. The search below does what string matching does: it compares code points.

A simplified demonstration of a documented mechanism, not a Zawgyi emulator. The encodings are labelled rather than rendered, because in a modern browser the two look alike, which is the difficulty.

The form nobody could read

In August 2018 Facebook stated that over 90 per cent of phones in Myanmar used Zawgyi. Its own Help Centre pages and reporting tools were in Unicode.

So the sharper failure is not the classifier. It is that the people in the best position to report incitement, in the language it was written in, could not read the form that would have let them report it.

Facebook removed Zawgyi as an interface option for new users in 2018 and completed automatic conversion in Facebook and Messenger in September 2019, roughly two years after the clearance operations in Rakhine State.

The staffing, as recorded

Mid 2014
One Burmese-speaking content moderator, contracted, based in Dublin. Myanmar users then numbered in the hundreds of thousands.
Early 2015
Reuters found two people at Facebook who could speak Burmese reviewing problematic posts. Before that, most reviewers of Burmese content spoke English.
2018
In Sri Lanka, after anti-Muslim riots, Facebook was reported to have had two resource persons reviewing Sinhala. A local organisation had first raised platform misuse with the company in 2009.
June 2018
Facebook stated it had over 60 Myanmar-language experts reviewing content, against roughly 20 million users, with a target of at least 100 by year end.
May 2021
An internal document, later disclosed: "We don't have coverage for Ethiopia due to lack of human review capacity there." Ethiopia had no hate speech classifier during an active civil conflict.
June 2022
Global Witness submitted twelve Amharic hate speech examples as paid advertisements, drawn from content already reported as violating. All twelve were approved. Two more, a week later, were approved within hours.

Two families, one state

Burmese is Sino-Tibetan. Rakhine is a Burmese lect, so also Sino-Tibetan. Rohingya is Indo-Aryan, at the far end of the line from the first module. The people who share Rakhine State share a territory, a history and a war, and no language family at all.

Which means the obvious remedy would not have worked even in principle. Meta's Myanmar failure was a failure in the perpetrators' language; the victims' language has never been offered an interface, a classifier or a translation engine by any platform anywhere.

The Hate Speech Dataset Catalogue lists Albanian, Arabic, Bengali, Chinese, Croatian, Danish, Dutch, English, Estonian, French, German, Greek, Hindi, Indonesian, Italian, Korean, Latvian, Polish, Portuguese, Russian, Slovene, Spanish, Turkish, Ukrainian and Urdu.

It lists neither Rohingya nor Burmese. The most notorious case of platform-enabled ethnic violence of the last decade produced no annotated dataset in the victims' language and none in the perpetrators'.

So the resourcing was not there. The last module is about why nothing required it to be.

Module 03

Where the liability stops

Myth we will unpick: platforms under-moderate because moderation is hard.

It is hard. It is also, in the jurisdiction where the decisions are taken, almost entirely free not to do. Two subsections explain most of that.

Section 230, in two halves

Enacted in 1996 as part of the Communications Decency Act. Choose a subsection.

The Rohingya claim, and how it ended

A class action was filed against Meta in California in December 2021, seeking more than 150 billion US dollars on behalf of Rohingya refugees. It was dismissed, and the dismissal was appealed. On 28 April 2026 the Ninth Circuit affirmed, holding that Section 230 barred every claim.

The plaintiffs had argued in the alternative that Myanmar law should apply. Before you read on: on what ground would you expect an American court to decline to apply it?

The court held that Myanmar's interest was "insufficiently incorporated into the positive law of the country", because no Myanmar case law supports publisher liability for social media companies.

Read plainly: a state whose legal system has collapsed cannot generate the jurisprudence that would give it a cognisable interest in a case about the platform implicated in the collapse.

The same opinion held that "matching users with content is publishing conduct, even when the user has not requested the content". Note also that the Supreme Court has never decided whether Section 230 covers recommendation algorithms; Gonzalez v. Google in 2023 is widely described as settling that and in fact declined to reach it.

Where liability does exist

Kenya's High Court rejected Meta's jurisdictional challenge on 3 April 2025, in the case brought over the killing of Meareg Amare Abrha, a Tigrayan professor shot outside his home in November 2021 after Facebook posts carrying his name, photograph and work address. The posts were removed eight days after his death, having been reported before it.

In a parallel case, in April 2026, the Employment and Labour Relations Court ordered a Meta director to be cross-examined on an affidavit filed for the company. Kenyan courts have rejected Meta's jurisdictional argument at every instance, on the ground that it derives substantial advertising revenue from the country.

The American front closed in April 2026. The Kenyan one did not.

One clause, twenty-four languages

Article 42 of the EU Digital Services Act requires very large platforms to publish, twice a year, the human resources they devote to content moderation broken down by each applicable official language of the Member States, together with those staff's linguistic qualifications. It is the only such obligation anywhere.

Covered by Article 42
Languages in this deck

The United Kingdom's Online Safety Act has no equivalent provision. An American attempt to create one, the LISTOS Act, was reintroduced in December 2025, which means earlier versions died.

Why this matters

The technical difficulty is real and it is not the binding constraint. The tell is Spanish: a language with abundant resources and an enormous American speaker population, whose posts are reportedly flagged at around half the English rate. If enforcement halves for Spanish, scarcity is not the explanation. Allocation is.

Of the misinformation classification budget disclosed in 2021, 87 per cent went to the United States. In the six months after Community Notes replaced fact-checking there, 900 notes were published in the US, against roughly 35 million posts labelled by professional fact-checkers in the EU. Community Notes supports six languages. Neither Bengali nor Burmese is among them.

There is no tidy recommendation at the end of this. The jurisdiction that compels per-language disclosure is the one whose languages were never where the violence happened, and the jurisdiction where the decisions are made has arranged its law so that the question does not arise. That is the position, and it is worth looking at squarely rather than resolving.

UN Independent International Fact-Finding Mission on Myanmar, A/HRC/39/64, 24 August 2018, paragraph 74. Facebook Engineering, "Removing Zawgyi", 26 September 2019; Facebook Newsroom, "An Update on Myanmar", 15 August 2018. BSR, Human Rights Impact Assessment: Facebook in Myanmar, October 2018. Amnesty International, The Social Atrocity, September 2022. Reuters special report on Facebook in Myanmar, 15 August 2018. Al Jazeera on Facebook's Sri Lanka apology, 13 May 2020. Rest of World, "Why Facebook keeps failing in Ethiopia", 13 November 2021. Global Witness, Ethiopia advertisement test, 9 June 2022. Amnesty International, A death sentence for my father, October 2023, and the Kenyan High Court jurisdiction ruling, 3 April 2025. United States Court of Appeals for the Ninth Circuit, Doe 1 v. Meta Platforms, Inc., No. 24-1672, 28 April 2026. Congressional Research Service R46751 on Section 230; Regulation (EU) 2022/2065, Article 42. Oversight Board, Assessing Meta's Plans to Expand Community Notes, 26 March 2026. Nieman Lab coverage, same date. Grierson, G. A., Linguistic Survey of India, 1903. Glottolog, on the 2007 retirement of ISO 639-3 cit. Translators without Borders, Cox's Bazar comprehension study, published 12 December 2018. Joshi, P., et al., "The State and Fate of Linguistic Diversity and Inclusion in the NLP World", ACL 2020. Hate Speech Dataset Catalogue, hatespeechdata.com. Per-claim confidence ratings, retrieval dates and unresolved items are held in the research files at the repository root.
Arrow keys, space, or click to move