Repository navigation
laya-multilingual never selects the first-listed score option — in English too (0/290), while laya does #131
Description
Activity
Follow-up, one more condition that separates position from label within a single run.
Condition F: per-item random option order (fixed seed, 300 Japanese items). Within any single fixed ordering, "rejects slot 1" and "rejects whichever word is in slot 1" predict the same table; shuffling per item breaks that.
observed if uniform first slot chosen 0 / 300 100 argmax by slot [0, 149, 151] [100, 100, 100] picks by label (not urgent / soon / blocked) 75 / 93 / 132 — times each label was placed in slot 1 90 / 109 / 101 100 each All three labels get chosen, and each was placed first about equally often, so every label lost exactly the picks it had while sitting in slot 1. Only slot 1 has a hole; slots 2 and 3 split evenly, so it is not a last-slot preference either. Accuracy 0.350.
Data and script:
docs/baseline_ja.md§6.2b in the repo above.I ran this down against the code and the weights, and the evidence points squarely at the
laya-multilingualcheckpoint's score head, not the inference code. Reproduced on CPU (deterministic, fp32), reading the raw marker logits straight offmodel(...)— before temperature and before softmax — so the decode arithmetic can't be a factor.1. Same code path, swap only the checkpoint → the hole appears
48 score decisions (8 score questions × 6 English states), identical code for both checkpoints:
checkpoint slot 0 chosen mean raw logit per slot (k=3 Qs) laya(English)12/48 = 25.0% slot0 +0.512, slot1 +1.474, slot2 −0.05 laya-multilingual4/48 = 8.3% slot0 −1.719, slot1 +0.111, slot2 −0.462 agent.py/common.pycontain no checkpoint-conditional branch — nothing in the code path knows which weights are loaded. A suppression that appears only on the multilingual weights therefore cannot originate in the code. On the English checkpoint slot 0 is healthy (+0.512); on multilingual it sits ~1.8 nats below slot 1.2. It's at the model output, not the decode
The numbers above are
logits[marker_pos]read directly from the forward pass, pre-temperature and pre-softmax. Slot-0's raw logit is −1.719 (multilingual) vs +0.512 (English). Soexp_score = (arange(k)*p).sum(), the temperature scaling, and the marker gather are all downstream of the defect and can't be causing it. (Thechoice:11+temperature-clamp warning is for choice questions and is irrelevant here — and moot for raw logits.)3. Control: identical option text isolates position from content
Three identical options (
["moderate","moderate","moderate"]), so the input differs by slot position only. Raw logits minus their mean:laya(English):[−0.075, +0.053, +0.022]— essentially flat.laya-multilingual:[−0.100, +0.507, −0.407]— a strong learned pattern (slot 1 boosted, slots 0 and 2 depressed).
A position-neutral input producing a sharply position-dependent output on multilingual, but a flat one on English, is the signature of a positional prior baked into the multilingual weights.
Conclusion
This is a training/weights artifact of
laya-multilingual's score head (a learned bias against slot 0), consistent with the invertedboolAUROC (0.355) you'd also see on that checkpoint — a pure decode/marker bug would not invert the bool head too. It is real and reproducible, but not fixable incommon.py/agent.py; it needs a retrain/rebalance of the multilingual checkpoint. Mitigation until then: on the multilingual checkpoint, avoid placing the most-likely level in slot 0, or add a neutral/sentinel level 0.(My multilingual run gave 8.3% rather than a strict 0% — my question set is milder than yours — but the raw-logit suppression is unambiguous and in the same direction.)
@hiroki-abe-58 nice — Condition F is the clean position-vs-label separator, and I reproduced it (below). I then ran one more cut that separates position from the
level N:text thatrender_optionsalways emits forscore, and it localizes the bias a bit further.I built score items with custom option strings (bypassing
render_options), reading raw marker logits on the multilingual checkpoint, averaged over 8 states, one 3-level question:option rendering mean raw logit / slot argmax per slot A: level 0/1/2at pos 0/1/2 (normal)[−0.49, +0.95, −0.46] [0, 8, 0] B: level 2/1/0at pos 0/1/2 (ordinal text reversed)[−0.44, +0.86, −0.42] [0, 8, 0] C: no levelprefix at all[+0.33, +0.08, −0.41] [3, 4, 1] D: word ordinals zero/one/two[−0.12, +0.35, −0.22] [2, 5, 1] Two things fall out:
- It's positional, not the ordinal-number token. B moves the literal
level 0:text to the last slot, and the pattern is unchanged ([0,8,0]) — the hole stays in slot 1 (1-indexed). So it's not thelevel 0string. This matches your Condition F. - But the hole is entangled with the
level N:render scaffold. Dropping that prefix entirely (C) makes the slot-0 suppression vanish (raw logit −0.49 → +0.33, argmax 0→3/8); word ordinals (D) mostly recover it too.render_optionsrenders score as"level %d: ...", and the multilingual score head appears to have learned a slot-1 prior coupled to that exact format.
Caveat, so this isn't over-read: the C/D recovery is in the raw logits / argmax spread only. The head was trained on the
level N:format, so dropping it is off-distribution — the slot-0 recovery could be the learned prior being scrambled rather than accuracy improving. Whether a rendering change is a real mitigation or just noise needs your labelled set (an accuracy delta for A vs C/D on the multilingual checkpoint would settle it). It's a lead to test, not a confirmed fix.Net, consistent with what I posted above: this lives in the
laya-multilingualscore head (a positional/format-coupled prior), not incommon.py/agent.py— the English checkpoint through the identical code path does not have the hole. A proper fix is a retrain/rebalance; a rendering tweak is worth an A/B on your set as a cheaper interim.- It's positional, not the ordinal-number token. B moves the literal
@AlKor13
This is excellent — thank you for going down to the raw logits and for the identical-option control; that settles "weights, not code" cleanly.On the
level N:scaffold: I have the labelled sets, so I'll run A vs C vs D on both bench_ja (300) and bench_en (290) with laya-multilingual, and the same three on the English checkpoint as a control, reporting first-slot counts, accuracy and RPS. That should tell us whether dropping the prefix is a real mitigation or just scrambles the prior. I'll post the numbers here tomorrow (JST).@AlKor13 Ran A vs C vs D on both labelled sets. Short version: you were right to be cautious. The rendering matters far more than I expected, but no single replacement is a fix.
Setup. Same harness as before; condition A reproduces
agent.predict()to 4.6e-5 and the published numbers exactly (bench_ja acc 0.447, slot 0 chosen 0/300; bench_en 0/290). C = nolevel N:prefix, D = word ordinals.Raw logits. On laya-multilingual, with the prefix, slot 0 sits 3.4 logits (ja) / 5.2 (en) below the other slots. Without it: 0.4 / 0.06. It is not a uniform offset either — under C the lowest slot moves to slot 2, so the ordering changes, not just the gap.
Accuracy, paired McNemar on identical items:
bench change score acc items that flipped correctness p ja A → C 0.447 → 0.513 56.7% 0.145 (n.s.) ja A → D 0.447 → 0.570 38.3% 0.0007 en A → C 0.266 → 0.428 56.9% 0.0003 en A → D 0.266 → 0.293 19.3% 0.350 (n.s.) Two of four are significant, and they are different conditions: D wins on Japanese, C wins on English, each is null on the other. So the supportable claim is "the shipped rendering is not the best one for this checkpoint", not "dropping the prefix fixes it".
What I'd flag most: the flipped column. Under C, 57% of items change correctness, far more than the 6.7 / 16.2 point accuracy delta. A cosmetic change to the option string very nearly re-rolls the prediction. "Suppresses slot 0" undersells what the prefix is doing.
Also: on bench_ja, D leaves slot 0 mostly suppressed (22/300, mean logit −2.9) yet gives the best accuracy and RPS. Slot-0 recovery and accuracy are not the same axis.
Control. The English
layacheckpoint under the same condition A shows no suppression (slot-0 logit +0.39 ja / +1.35 en), and removing the prefix makes it worse on bench_en (0.583 → 0.500). Solevel N:is not harmful in general; laya-multilingual has learned something wrong about that specific pattern. Consistent with your "weights, not code".Net: a rendering tweak is not a safe interim mitigation for this checkpoint — it trades one instability for another. Retrain / rebalance remains the fix. Full tables incl. RPS and per-slot logits:
docs/baseline_ja.mdanddocs/baseline_en.md§6.2c, raw data inruns/option_rendering*.jsonin the repo.Thank you both. This is a model of how to run a bug down: Condition F to separate position from label, the raw logits to rule out the decode, and the identical-option control to rule out content.
Where this leaves it, as I read your results:
- It is in the weights, not the code. Same code path, and only
laya-multilingualshows the slot-0 hole. The English checkpoint under the same condition does not. - A rendering change is not a safe interim fix. @hiroki-abe-58's A/C/D run shows 57% of items flipping correctness under a cosmetic change, with the better variant differing by language. Shipping a new default render on that basis would swap one instability for another, so I'm not going to.
- The real fix is in training: ordinal options need the same position-balancing the choice questions get, so the score head cannot learn a slot prior.
In the meantime I'll add this to the README's Limits section:
laya-multilingualhas a measured position prior onscorequestions, so for English score questions use the English checkpoint (model="english"), and for other languages validate score outputs on your own data before relying on them.- It is in the weights, not the code. Same code path, and only
0.3.7, released today, documents this in the README's Limits section:
laya-multilingualhas a measured position bias onscorequestions. For English score questions, route tomodel="english"; for other languages, validate score outputs on your own data. As you both showed, the actual fix is in training, and this issue stays open until a retrained checkpoint lands. Thank you again for the analysis.@NandhaKishorM
Thanks for shipping the Limits note in 0.3.7 so quickly, and for keeping this open until the retrain.When a position-balanced multilingual checkpoint is ready, I'm happy to run the same before/after on it — conditions A–F plus the A/C/D rendering check on bench_ja (300) and bench_en (290) — and post the table here. The benches are CC BY 4.0 in the repo if you'd rather run them yourself.
Thank you @hiroki-abe-58, I'll take you up on that. When a position-balanced multilingual checkpoint is ready, I'll ping you here so the same A to F conditions on bench_ja and bench_en can serve as the before/after.
@NandhaKishorM Thanks — I'll be here. In the meantime I opened #259, a label-free check for this that you can run on the retrained checkpoint yourself (current multilingual fails it, english passes); it builds on @AlKor13's identical-option control above. The bench_ja / bench_en A–F before/after stays on offer on top of that.
- added a commit that references this issue
on Sep 23, 2026 Thank you @hiroki-abe-58. #259 is merged, so the label-free check is ready for the retrained checkpoint.
Reacted by GeneLab- added 2 commits that reference this issue
on Sep 24, 2026 There is now a structural answer to this, though not a published checkpoint yet, so I am leaving this open.
#951 merged in 0.3.28 and adds an opt-in parallel option layout: every option starts at the same position id, an option cannot attend to another option's tokens, and ModernBERT's local layers measure their sliding window in position ids rather than sequence index. The decision head has no positional encoding, so with order-blind marker embeddings the logits simply permute with the options. Reordering can only reorder the answer.
Measured on typed-decisions test, 400 cases and 2,000 decisions, three arms fine-tuned the same way:
arm accuracy, listed order accuracy, shuffled same answer when shuffled max probability change sequential, current 0.779 0.767 91.7% 0.267 sequential plus shuffle_options0.757 0.750 95.0% 0.257 parallel 0.764 0.764 100% 0.003 The 12.6% figure in that PR is the
scoreflip rate on the sequential arm, which is the question type this issue is about, and it goes to zero.Why this does not close yet. The layout is a property of the weights, so it lives in the checkpoint config, and every published checkpoint stays on the default sequential layout. Switching the shipped weights over without retraining is exactly invariant and much worse: MASSIVE accuracy falls from 0.795 to 0.54, because the weights learned the other layout. So fixing this issue means training and publishing a parallel multilingual checkpoint, which is mine to do.
What the PR gives me is the evidence to justify that compute, which I did not have before. The honest caveats from it, which apply to your case: one seed per arm, differences under about 0.01 should not be read as real, the parallel arm trails slightly on
choice, and the layout removes the effect of option position, not of option text. Thelevel N:prefix staying with its slot meanspresentation_checks.pystill shows a non-zero effect there, so if part of what you observed is the model reading the level labels rather than their positions, this will reduce it rather than remove it.Also relevant and already shipped: a fine-tune with
shuffle_options=("choice",)took the answer-change rate from 14/90 to 0/90 under reorder on a small real-weights task (#899), which is the approximate version of the same fix and is available now.Leaving open against publishing a parallel checkpoint. Thank you for a report with a 0/290 count in it; a bias that absolute is what made it worth taking structurally rather than tuning around.
Thank you, this is a much better answer than tuning around it, and the 100% / 0.003 row is exactly what the report hoped for.
Two things I can take on, whenever useful:
- When a parallel multilingual checkpoint exists, I'll run presentation_checks on it in ja/ko/hi/tr plus bench_ja / bench_en and post the numbers here.
- On the level N: caveat: in Lev's review (Measure Score's position effect, and offer order averaging for Score (off by default) Abhinavexists/lev#2) we found numbered keys carry an order of their own, and replaced them with unordered shape keys averaged over every key-to-slot assignment. A Score variant of the identical control built that way would separate "reads the position" from "reads the level label", on both layouts. Happy to send it as a small PR if you'd like that split measured.
I initially found this in Japanese and assumed it was a language gap, so I re-ran it in English. It is not a language gap.
Setup: 300 Japanese / 290 English label-conditioned business emails, three questions each (4-way
choice, 3-level ordinalscore, onebool). Inference vialaya.load()/agent.predict()exactly as in the README (laya 0.3.4); my harness reproduces the package's probabilities to float precision.score, first-listed option chosen, five conditions (original / reversed / relabeled Low-Mid-High / relabeled reversed / 4 levels):laya-multilinguallaya(English ckpt)For
laya-multilingualin English, "Not urgent" is chosen 0 times when listed first and 285 times when listed last; the mass follows the slot, not the label. The English checkpoint picks the first option 22–26% of the time under the same conditions, so it is not the task or the harness.In English,
laya-multilingualalso getsscoreRPS 0.340 (random 0.197) andboolAUROC 0.355 (below 0.5 = inverted ranking). English is a training language for that checkpoint, so "can't read the input" does not explain it. It looks like something specific to the multilingual checkpoint — possibly the marker layout, possibly the ordinal training.Full data, raw probabilities and repro scripts: https://github.com/hiroki-abe-58/sokudan/blob/main/docs/baseline_en.md (and
baseline_ja.md). Happy to help debug.