Blog··i18n

Arabic is not a translation pass

A product built in English and translated afterwards works in English and apologises in everything else. Whity is built in Jordan, where that is not an acceptable outcome — so Arabic shaped the schema, not just the strings.

en · LTRar · RTLforms.incidentIncident reportDepartment headبلاغ حادثةرئيس القسمone key · two directions · neither is the fallback

You can tell when a second language arrived late. Labels are English keys with an Arabic gloss bolted on somewhere else. A name field holds one string, so the Arabic version lives in a parallel table nobody validates. The layout flips with a stylesheet and something important ends up on the wrong side of the screen. None of it is broken enough to fail a test, and all of it is obvious to the person reading it in Arabic.

Treating Arabic as a first-class case costs more up front, in four specific places.

1. Localized text is a type, not a convention

In Whity's API schema, a human-facing name is not a string. It is an object carrying both languages, marked with x-whity-localized-text so every generated client and every consumer knows what it is looking at:

POST /api/v1/formsrequest schema
{
  "form_key": "incident-report",
  "name": {
    "ar": "بلاغ حادثة",
    "en": "Incident report"
  }
}
// not a string with a translation table somewhere else.
// both languages travel together, or neither does.

This is the decision everything else rests on. When the field itself carries both languages, a form created through the API arrives complete, an agent calling the same route supplies both, and there is no code path that can save a record in one language and lose the other — because there is no separate place for the other to live.

2. English is generated from the code that renders it

There is no en.json anybody edits. Every user-facing string goes through a translation call whose second argument is the English text:

a screenTSX
const t = useTranslation('auth');

<Button>{t('login.submit', 'Sign in')}</Button>

An extract command derives the English catalogue from those call sites, and CI fails if the committed catalogue has drifted from the source. Three consequences follow, and they are the reason for the whole arrangement:

  • the screen still reads normally in a diff — you see the English where it renders;
  • it renders correctly before any translation bundle arrives, or if a key was never seeded;
  • English cannot drift from the code, because it is the code. There is no second list of strings for someone to forget.

It also means an extraction effort can fan out across many people without two of them editing the same file, which is the practical reason a large codebase ever finishes this work.

3. The sync adds. It never overwrites.

Here is the failure everyone building this eventually has: a translator spends a week filling in Arabic, someone runs the sync, and it is gone — with nothing in a diff to notice, because the database is not in the diff.

Whity's catalogue is a mirror of the code: a key that leaves the source leaves the file, and CI insists on it. The database is an accumulator: it only ever gains rows, because what is in it is human work. Nothing flows backwards.

t('key', 'English')the sourceextracti18n/<domain>.jsongenerated · CI-checkedsynctranslations tableinsert-onlya humanwrites arnothing flows backwards · no UPDATE, no DELETE, anywhere in the sync
The one-way pipeline. The catalogue mirrors the code; the table accumulates human work.

The guarantee is structural rather than careful. The sync class containsno UPDATE and no DELETE statement at all, and a unit test asserts that about the class's own source text. You are not trusting aWHERE clause to be right; you are trusting that a statement which does not exist cannot run.

It also never machine-translates and never copies English into another language's rows. That sounds like a missing feature until you consider what the alternative produces: a row containing English text while claiming to be Arabic is indistinguishable from a finished translation — to the coverage report, to the console, and to the reviewer deciding what is left to do. An untranslated key has to be visibly untranslated, which means having no row at all.

4. Layouts mirror. They do not flip.

RTL is not text-align: right. A mirrored interface moves the whole axis: the sidebar, the back button, the order of a breadcrumb, which side a border sits on, which way a progress arrow points. Whity's design system is built on CSS logical properties —margin-inline, border-inline-start, text-align: start— so a screen mirrors as a unit rather than needing a second stylesheet that someone forgets to update.

Typography is part of it: the token package ships Noto Sans Arabic alongside Noto Sans, so Arabic renders in a face chosen for it rather than whatever the operating system substitutes. In the document designer, text direction is a per-block setting — a bilingual contract can carry an Arabic clause beside its English rendering on the same page, which a global toggle cannot express.

And languages themselves are data. There is an API for them — create a language, list them, report coverage per domain — rather than a constant in the source. Adding a third language is an operation, not a release.

The test that matters is boring. Open any screen with the direction reversed and read it. Not "does it render" — does the eye land in the right place, does the primary action sit where a hand expects it, does a sentence containing an English product name inside an Arabic paragraph break correctly. Bidirectional text has opinions, and they only show up when you look.

What this does not solve

Plenty. Arabic pluralisation is richer than the two-form case most i18n tooling assumes. Sorting names in a language with different collation rules is its own project. PDF rendering needs the right fonts present in the render container, which is an infrastructure concern that bites the first time you generate a document on a fresh host. None of that is fixed by the four decisions above — they just mean the foundation is not fighting you while you deal with it.

The ambition here is modest and specific: Arabic should not be the case that got tested last. Everything above exists so that the answer to "does it work in Arabic" is decided by the schema and the layout system, rather than by whoever remembered.

Frequently asked

Does Whity support Arabic and right-to-left layouts?

Yes, as a first-class case rather than a translation layer. Localized text is a type in the API schema carrying ar and en values side by side, the design system ships Noto Sans Arabic, layouts use CSS logical properties so screens mirror rather than flip, and the document designer treats text direction as a per-block setting.

How are translations managed in Whity?

The English catalogue is generated from the code that renders it — every string goes through t('key', 'English text'), and an extract command derives the catalogue from those call sites, with CI failing if it drifts. A sync command loads catalogues into the translations table, and translators fill other languages through an admin screen. Languages themselves are data, managed through the API, not a hardcoded list.

Will running a translation sync overwrite work translators have done?

No. The sync class contains no UPDATE and no DELETE statement at all, and a unit test asserts that about the class's own source text — so the guarantee is structural rather than a promise about a WHERE clause. Existing rows are never written to in any language or scope; the sync only ever adds.

Does Whity machine-translate strings into Arabic?

No. The sync never machine-translates and never copies English into another language's rows. A row containing English text while claiming to be Arabic is indistinguishable from a finished translation, so an untranslated key is represented by having no row at all — visibly missing rather than quietly wrong.

The extraction pipeline, the CI gate and the sync guarantees are documented in full inthe internationalization guide.