The Translation Lab
Baudelaire Goes to the Model Arena: Literary Translation
Nine AI translation pipelines, four languages, one magnificently unpleasant poem—and a Silver model that beat Gold.

In the previous Field Note, we discussed a rather provocative paper showing that readers could not reliably distinguish AI-generated poetry from poetry written by humans and, even more inconveniently, sometimes preferred the machine-made poems. That naturally left us with another question: if today's LLMs can produce something that looks and feels like poetry, can our translation pipeline preserve poetry when moving it from one language to another? There was only one reasonable way to find out. We went back to Baudelaire.
There are reasonable texts to use when testing a translation system. Contracts are reasonable. Instruction manuals are reasonable. Newspaper articles are reasonable. They contain sentences that generally want to be understood, which is a useful quality when you are trying to measure whether a translation works.
We chose Baudelaire.
More precisely, Spleen (LXXVIII) from Les Fleurs du mal, a magnificently unpleasant poem in which the sky becomes a lid, the earth turns into a wet prison, Hope flaps around like a bat against rotten ceilings, spiders install themselves inside our brains and Despair eventually plants a black flag on the poet's skull.
It seemed perfect.
Not because we expected to settle whether AI can “translate poetry”, a question large enough to keep translators and moderately drunk philosophers occupied indefinitely, but because we wanted to ask something smaller: given the models available today, what happens when we build a literary translation pipeline around them, and how much do we have to spend before the translation actually gets better?
Why Baudelaire?
Spleen is a cruel test because most of the things that make it work are exactly the things that disappear when translation becomes sentence replacement. The poem accumulates rather than merely proceeds: stanza after stanza begins with “Quand”, the syntax keeps stretching forward and the images become progressively more suffocating, until the reader is trapped inside the sentence with Baudelaire.
A translation can get every noun right and still kill the poem.
That is precisely the sort of problem we care about at TranswrAIte. Semantic accuracy is necessary, but literary translation lives in a more troublesome neighbourhood, where voice, rhythm, repetition and even punctuation matter. Sometimes a sentence must remain strange because the original is strange, and polishing it into impeccable contemporary prose is precisely the mistake.
Baudelaire gives a translation system plenty of rope with which to hang itself. We accepted the offer.
Nine pipelines walk into a poem
We tested OpenAI, Anthropic and DeepSeek pipelines on the same French poem, translating it into Italian, English, Spanish and German, with configurations we called Bronze, Silver and Gold.
Every translation required two calls rather than one. A first model produced the draft, then an editorial model compared it with the French original and corrected omissions, mistranslations, awkward phrasing, punctuation, formatting and failures of voice. Only this reviewed version entered the benchmark.
This is much closer to how we think LLM literary translation should work: not one gigantic artificial translator confronting Baudelaire alone, but something resembling a tiny editorial office where somebody drafts, somebody rereads and somebody quietly changes the sentence.
The LLM version has the advantage that nobody needs coffee.
Then came the judges
The finished translations were anonymously evaluated by OpenAI GPT-5.6-sol and Anthropic Claude Opus 5, with fidelity, voice and naturalness carrying most of the score. We also deliberately damaged translations by removing passages, duplicating sentences and changing numbers; both judge families detected 100% of the stored planted defects.
This does not transform an LLM into a great literary critic. Detecting a missing paragraph and deciding whether Baudelaire has survived his journey into German are very different intellectual activities, but at least we knew the judges could recognise a corpse when one was lying on the floor.
Silver beat Gold
Then our expensive horse lost.
OpenAI Silver achieved the highest average quality at 4.129/5, followed by OpenAI Gold at 4.050 and Anthropic Gold at 3.879. The interesting part appears beside the bill: Silver cost roughly $0.057 per completed translation and review, while Gold cost about $0.226.
Gold was approximately four times more expensive and slightly worse.
There is a particular pleasure in a benchmark knocking your tidy assumptions off the table. We had built Gold to be Gold: more reasoning, more expensive inference, larger budgets. You expect it to enter wearing the better suit. Instead Silver came through the kitchen and stole its chair.
The quality-versus-price graph makes this particularly clear. Silver sits near the top at around six cents, while Gold has travelled far to the right without climbing any higher; Anthropic Bronze, meanwhile, reaches a respectable 3.858/5 for roughly $0.019.

For one poem these differences are pocket change. Across novels, catalogues and multiple languages, they become architecture.
There was no “best model”
Things became even less tidy when we looked at individual languages. Anthropic Gold won Italian and Spanish, OpenAI Silver won English, and OpenAI Bronze won German.
So the overall winner did not sweep the languages, and one Bronze configuration beat everybody in German.
This suggests that “Which model is best for translation?” may simply be the wrong question. Italian Baudelaire and German Baudelaire are already different problems, before we even reach dialogue, historical prose, dialect or a 500-page novel whose narrator changes register halfway through.
The more useful question is which model is best for which language, at which stage, for which kind of text?
Once you begin thinking that way, sending an entire book through one enormous model starts to look rather primitive.
And then we threw humans into the machine
We also imported one existing human translation for each language as an anonymous candidate. The judges were not told they were human, because calling something a “human reference” gives it a little crown before the competition begins.
The results were peculiar: English scored 3.975, Italian 3.600, German 2.575 and Spanish only 1.750.
Our conclusion was emphatically not that AI had defeated human translators. The Spanish and German versions may be freer adaptations, structurally different editions or translations pursuing objectives our rubric penalises, and those results require manual examination rather than triumphal declarations about machines conquering literature.
Still, the experiment left us with a principle worth keeping: human is provenance, not a quality score, and AI is provenance too.
Eventually there is only a text on the page, and somebody has to read it.
What we learned
The interesting engineering problem is becoming less about persuading one enormous model to translate a book beautifully in a heroic single pass, and more about building a system capable of choosing the right model, reviewing its work, detecting suspicious passages and leaving alone the sentences that already work.
In other words, the future of AI literary translation may look less like one brilliant artificial translator and more like a slightly neurotic publishing house, with several editors leaning over the same manuscript and disagreeing about a comma.
This experiment was only one poem in four languages, so it proves very little about literature at large. But it gave us three useful clues: more expensive does not necessarily mean better, different languages reward different models, and the pipeline matters at least as much as the model sitting inside it.
For this particular rainy, claustrophobic, spider-infested trip through Baudelaire's skull, Silver won.
Gold cost four times more and came second, which feels, somehow, like an ending Baudelaire might have appreciated.
Spleen (LXXVIII)
Charles Baudelaire · Les Fleurs du mal
Quand le ciel bas et lourd pèse comme un couvercle
Sur l'esprit gémissant en proie aux longs ennuis,
Et que de l'horizon embrassant tout le cercle
Il nous verse un jour noir plus triste que les nuits ;
Quand la terre est changée en un cachot humide,
Où l'Espérance, comme une chauve-souris,
S'en va battant les murs de son aile timide
Et se cognant la tête à des plafonds pourris ;
Quand la pluie étalant ses immenses traînées
D'une vaste prison imite les barreaux,
Et qu'un peuple muet d'infâmes araignées
Vient tendre ses filets au fond de nos cerveaux,
Des cloches tout à coup sautent avec furie
Et lancent vers le ciel un affreux hurlement,
Ainsi que des esprits errants et sans patrie
Qui se mettent à geindre opiniâtrement.
— Et de longs corbillards, sans tambours ni musique,
Défilent lentement dans mon âme ; l'Espoir,
Vaincu, pleure, et l'Angoisse atroce, despotique,
Sur mon crâne incliné plante son drapeau noir.