Skip to content

laya-multilingual never selects the first-listed score option — in English too (0/290), while laya does #131

Description

@hiroki-abe-58

I initially found this in Japanese and assumed it was a language gap, so I re-ran it in English. It is not a language gap.

Setup: 300 Japanese / 290 English label-conditioned business emails, three questions each (4-way choice, 3-level ordinal score, one bool). Inference via laya.load() / agent.predict() exactly as in the README (laya 0.3.4); my harness reproduces the package's probabilities to float precision.

score, first-listed option chosen, five conditions (original / reversed / relabeled Low-Mid-High / relabeled reversed / 4 levels):

model Japanese (n=300) English (n=290)
laya-multilingual 0, 0, 1, 1, 0 0, 0, 0, 0, 0
laya (English ckpt) 13, 56, 8, 1, 110 65, 74, 0, 5, 4

For laya-multilingual in English, "Not urgent" is chosen 0 times when listed first and 285 times when listed last; the mass follows the slot, not the label. The English checkpoint picks the first option 22–26% of the time under the same conditions, so it is not the task or the harness.

In English, laya-multilingual also gets score RPS 0.340 (random 0.197) and bool AUROC 0.355 (below 0.5 = inverted ranking). English is a training language for that checkpoint, so "can't read the input" does not explain it. It looks like something specific to the multilingual checkpoint — possibly the marker layout, possibly the ordinal training.

Full data, raw probabilities and repro scripts: https://github.com/hiroki-abe-58/sokudan/blob/main/docs/baseline_en.md (and baseline_ja.md). Happy to help debug.

Activity

  1. hiroki-abe-58 commented on Sep 22, 2026

    @hiroki-abe-58
    ContributorAuthor

    Follow-up, one more condition that separates position from label within a single run.

    Condition F: per-item random option order (fixed seed, 300 Japanese items). Within any single fixed ordering, "rejects slot 1" and "rejects whichever word is in slot 1" predict the same table; shuffling per item breaks that.

    observed if uniform
    first slot chosen 0 / 300 100
    argmax by slot [0, 149, 151] [100, 100, 100]
    picks by label (not urgent / soon / blocked) 75 / 93 / 132 —
    times each label was placed in slot 1 90 / 109 / 101 100 each

    All three labels get chosen, and each was placed first about equally often, so every label lost exactly the picks it had while sitting in slot 1. Only slot 1 has a hole; slots 2 and 3 split evenly, so it is not a last-slot preference either. Accuracy 0.350.

    Data and script: docs/baseline_ja.md §6.2b in the repo above.

  2. AlKor13 commented on Sep 22, 2026

    @AlKor13
    Contributor

    I ran this down against the code and the weights, and the evidence points squarely at the laya-multilingual checkpoint's score head, not the inference code. Reproduced on CPU (deterministic, fp32), reading the raw marker logits straight off model(...) — before temperature and before softmax — so the decode arithmetic can't be a factor.

    1. Same code path, swap only the checkpoint → the hole appears

    48 score decisions (8 score questions × 6 English states), identical code for both checkpoints:

    checkpoint slot 0 chosen mean raw logit per slot (k=3 Qs)
    laya (English) 12/48 = 25.0% slot0 +0.512, slot1 +1.474, slot2 −0.05
    laya-multilingual 4/48 = 8.3% slot0 −1.719, slot1 +0.111, slot2 −0.462

    agent.py/common.py contain no checkpoint-conditional branch — nothing in the code path knows which weights are loaded. A suppression that appears only on the multilingual weights therefore cannot originate in the code. On the English checkpoint slot 0 is healthy (+0.512); on multilingual it sits ~1.8 nats below slot 1.

    2. It's at the model output, not the decode

    The numbers above are logits[marker_pos] read directly from the forward pass, pre-temperature and pre-softmax. Slot-0's raw logit is −1.719 (multilingual) vs +0.512 (English). So exp_score = (arange(k)*p).sum(), the temperature scaling, and the marker gather are all downstream of the defect and can't be causing it. (The choice:11+ temperature-clamp warning is for choice questions and is irrelevant here — and moot for raw logits.)

    3. Control: identical option text isolates position from content

    Three identical options (["moderate","moderate","moderate"]), so the input differs by slot position only. Raw logits minus their mean:

    • laya (English): [−0.075, +0.053, +0.022] — essentially flat.
    • laya-multilingual: [−0.100, +0.507, −0.407] — a strong learned pattern (slot 1 boosted, slots 0 and 2 depressed).

    A position-neutral input producing a sharply position-dependent output on multilingual, but a flat one on English, is the signature of a positional prior baked into the multilingual weights.

    Conclusion

    This is a training/weights artifact of laya-multilingual's score head (a learned bias against slot 0), consistent with the inverted bool AUROC (0.355) you'd also see on that checkpoint — a pure decode/marker bug would not invert the bool head too. It is real and reproducible, but not fixable in common.py/agent.py; it needs a retrain/rebalance of the multilingual checkpoint. Mitigation until then: on the multilingual checkpoint, avoid placing the most-likely level in slot 0, or add a neutral/sentinel level 0.

    (My multilingual run gave 8.3% rather than a strict 0% — my question set is milder than yours — but the raw-logit suppression is unambiguous and in the same direction.)

  3. AlKor13 commented on Sep 22, 2026

    @AlKor13
    Contributor

    @hiroki-abe-58 nice — Condition F is the clean position-vs-label separator, and I reproduced it (below). I then ran one more cut that separates position from the level N: text that render_options always emits for score, and it localizes the bias a bit further.

    I built score items with custom option strings (bypassing render_options), reading raw marker logits on the multilingual checkpoint, averaged over 8 states, one 3-level question:

    option rendering mean raw logit / slot argmax per slot
    A: level 0/1/2 at pos 0/1/2 (normal) [−0.49, +0.95, −0.46] [0, 8, 0]
    B: level 2/1/0 at pos 0/1/2 (ordinal text reversed) [−0.44, +0.86, −0.42] [0, 8, 0]
    C: no level prefix at all [+0.33, +0.08, −0.41] [3, 4, 1]
    D: word ordinals zero/one/two [−0.12, +0.35, −0.22] [2, 5, 1]

    Two things fall out:

    1. It's positional, not the ordinal-number token. B moves the literal level 0: text to the last slot, and the pattern is unchanged ([0,8,0]) — the hole stays in slot 1 (1-indexed). So it's not the level 0 string. This matches your Condition F.
    2. But the hole is entangled with the level N: render scaffold. Dropping that prefix entirely (C) makes the slot-0 suppression vanish (raw logit −0.49 → +0.33, argmax 0→3/8); word ordinals (D) mostly recover it too. render_options renders score as "level %d: ...", and the multilingual score head appears to have learned a slot-1 prior coupled to that exact format.

    Caveat, so this isn't over-read: the C/D recovery is in the raw logits / argmax spread only. The head was trained on the level N: format, so dropping it is off-distribution — the slot-0 recovery could be the learned prior being scrambled rather than accuracy improving. Whether a rendering change is a real mitigation or just noise needs your labelled set (an accuracy delta for A vs C/D on the multilingual checkpoint would settle it). It's a lead to test, not a confirmed fix.

    Net, consistent with what I posted above: this lives in the laya-multilingual score head (a positional/format-coupled prior), not in common.py/agent.py — the English checkpoint through the identical code path does not have the hole. A proper fix is a retrain/rebalance; a rendering tweak is worth an A/B on your set as a cheaper interim.

  4. hiroki-abe-58 commented on Sep 22, 2026

    @hiroki-abe-58
    ContributorAuthor

    @AlKor13
    This is excellent — thank you for going down to the raw logits and for the identical-option control; that settles "weights, not code" cleanly.

    On the level N: scaffold: I have the labelled sets, so I'll run A vs C vs D on both bench_ja (300) and bench_en (290) with laya-multilingual, and the same three on the English checkpoint as a control, reporting first-slot counts, accuracy and RPS. That should tell us whether dropping the prefix is a real mitigation or just scrambles the prior. I'll post the numbers here tomorrow (JST).

  5. hiroki-abe-58 commented on Sep 22, 2026

    @hiroki-abe-58
    ContributorAuthor

    @AlKor13 Ran A vs C vs D on both labelled sets. Short version: you were right to be cautious. The rendering matters far more than I expected, but no single replacement is a fix.

    Setup. Same harness as before; condition A reproduces agent.predict() to 4.6e-5 and the published numbers exactly (bench_ja acc 0.447, slot 0 chosen 0/300; bench_en 0/290). C = no level N: prefix, D = word ordinals.

    Raw logits. On laya-multilingual, with the prefix, slot 0 sits 3.4 logits (ja) / 5.2 (en) below the other slots. Without it: 0.4 / 0.06. It is not a uniform offset either — under C the lowest slot moves to slot 2, so the ordering changes, not just the gap.

    Accuracy, paired McNemar on identical items:

    bench change score acc items that flipped correctness p
    ja A → C 0.447 → 0.513 56.7% 0.145 (n.s.)
    ja A → D 0.447 → 0.570 38.3% 0.0007
    en A → C 0.266 → 0.428 56.9% 0.0003
    en A → D 0.266 → 0.293 19.3% 0.350 (n.s.)

    Two of four are significant, and they are different conditions: D wins on Japanese, C wins on English, each is null on the other. So the supportable claim is "the shipped rendering is not the best one for this checkpoint", not "dropping the prefix fixes it".

    What I'd flag most: the flipped column. Under C, 57% of items change correctness, far more than the 6.7 / 16.2 point accuracy delta. A cosmetic change to the option string very nearly re-rolls the prediction. "Suppresses slot 0" undersells what the prefix is doing.

    Also: on bench_ja, D leaves slot 0 mostly suppressed (22/300, mean logit −2.9) yet gives the best accuracy and RPS. Slot-0 recovery and accuracy are not the same axis.

    Control. The English laya checkpoint under the same condition A shows no suppression (slot-0 logit +0.39 ja / +1.35 en), and removing the prefix makes it worse on bench_en (0.583 → 0.500). So level N: is not harmful in general; laya-multilingual has learned something wrong about that specific pattern. Consistent with your "weights, not code".

    Net: a rendering tweak is not a safe interim mitigation for this checkpoint — it trades one instability for another. Retrain / rebalance remains the fix. Full tables incl. RPS and per-slot logits: docs/baseline_ja.md and docs/baseline_en.md §6.2c, raw data in runs/option_rendering*.json in the repo.

  6. NandhaKishorM commented on Sep 22, 2026

    @NandhaKishorM
    Owner

    Thank you both. This is a model of how to run a bug down: Condition F to separate position from label, the raw logits to rule out the decode, and the identical-option control to rule out content.

    Where this leaves it, as I read your results:

    • It is in the weights, not the code. Same code path, and only laya-multilingual shows the slot-0 hole. The English checkpoint under the same condition does not.
    • A rendering change is not a safe interim fix. @hiroki-abe-58's A/C/D run shows 57% of items flipping correctness under a cosmetic change, with the better variant differing by language. Shipping a new default render on that basis would swap one instability for another, so I'm not going to.
    • The real fix is in training: ordinal options need the same position-balancing the choice questions get, so the score head cannot learn a slot prior.

    In the meantime I'll add this to the README's Limits section: laya-multilingual has a measured position prior on score questions, so for English score questions use the English checkpoint (model="english"), and for other languages validate score outputs on your own data before relying on them.

  7. added a commit that references this issue on Sep 23, 2026
  8. NandhaKishorM commented on Sep 23, 2026

    @NandhaKishorM
    Owner

    0.3.7, released today, documents this in the README's Limits section: laya-multilingual has a measured position bias on score questions. For English score questions, route to model="english"; for other languages, validate score outputs on your own data. As you both showed, the actual fix is in training, and this issue stays open until a retrained checkpoint lands. Thank you again for the analysis.

  9. hiroki-abe-58 commented on Sep 23, 2026

    @hiroki-abe-58
    ContributorAuthor

    @NandhaKishorM
    Thanks for shipping the Limits note in 0.3.7 so quickly, and for keeping this open until the retrain.

    When a position-balanced multilingual checkpoint is ready, I'm happy to run the same before/after on it — conditions A–F plus the A/C/D rendering check on bench_ja (300) and bench_en (290) — and post the table here. The benches are CC BY 4.0 in the repo if you'd rather run them yourself.

  10. NandhaKishorM commented on Sep 23, 2026

    @NandhaKishorM
    Owner

    Thank you @hiroki-abe-58, I'll take you up on that. When a position-balanced multilingual checkpoint is ready, I'll ping you here so the same A to F conditions on bench_ja and bench_en can serve as the before/after.

  11. hiroki-abe-58 commented on Sep 23, 2026

    @hiroki-abe-58
    ContributorAuthor

    @NandhaKishorM Thanks — I'll be here. In the meantime I opened #259, a label-free check for this that you can run on the retrained checkpoint yourself (current multilingual fails it, english passes); it builds on @AlKor13's identical-option control above. The bench_ja / bench_en A–F before/after stays on offer on top of that.

  12. NandhaKishorM commented on Sep 23, 2026

    @NandhaKishorM
    Owner

    Thank you @hiroki-abe-58. #259 is merged, so the label-free check is ready for the retrained checkpoint.

  13. NandhaKishorM commented on Oct 5, 2026

    @NandhaKishorM
    Owner

    There is now a structural answer to this, though not a published checkpoint yet, so I am leaving this open.

    #951 merged in 0.3.28 and adds an opt-in parallel option layout: every option starts at the same position id, an option cannot attend to another option's tokens, and ModernBERT's local layers measure their sliding window in position ids rather than sequence index. The decision head has no positional encoding, so with order-blind marker embeddings the logits simply permute with the options. Reordering can only reorder the answer.

    Measured on typed-decisions test, 400 cases and 2,000 decisions, three arms fine-tuned the same way:

    arm accuracy, listed order accuracy, shuffled same answer when shuffled max probability change
    sequential, current 0.779 0.767 91.7% 0.267
    sequential plus shuffle_options 0.757 0.750 95.0% 0.257
    parallel 0.764 0.764 100% 0.003

    The 12.6% figure in that PR is the score flip rate on the sequential arm, which is the question type this issue is about, and it goes to zero.

    Why this does not close yet. The layout is a property of the weights, so it lives in the checkpoint config, and every published checkpoint stays on the default sequential layout. Switching the shipped weights over without retraining is exactly invariant and much worse: MASSIVE accuracy falls from 0.795 to 0.54, because the weights learned the other layout. So fixing this issue means training and publishing a parallel multilingual checkpoint, which is mine to do.

    What the PR gives me is the evidence to justify that compute, which I did not have before. The honest caveats from it, which apply to your case: one seed per arm, differences under about 0.01 should not be read as real, the parallel arm trails slightly on choice, and the layout removes the effect of option position, not of option text. The level N: prefix staying with its slot means presentation_checks.py still shows a non-zero effect there, so if part of what you observed is the model reading the level labels rather than their positions, this will reduce it rather than remove it.

    Also relevant and already shipped: a fine-tune with shuffle_options=("choice",) took the answer-change rate from 14/90 to 0/90 under reorder on a small real-weights task (#899), which is the approximate version of the same fix and is available now.

    Leaving open against publishing a parallel checkpoint. Thank you for a report with a 0/290 count in it; a bias that absolute is what made it worth taking structurally rather than tuning around.

  14. hiroki-abe-58 commented on Oct 5, 2026

    @hiroki-abe-58
    ContributorAuthor

    Thank you, this is a much better answer than tuning around it, and the 100% / 0.003 row is exactly what the report hoped for.

    Two things I can take on, whenever useful:

    1. When a parallel multilingual checkpoint exists, I'll run presentation_checks on it in ja/ko/hi/tr plus bench_ja / bench_en and post the numbers here.
    2. On the level N: caveat: in Lev's review (Measure Score's position effect, and offer order averaging for Score (off by default) Abhinavexists/lev#2) we found numbered keys carry an order of their own, and replaced them with unordered shape keys averaged over every key-to-slot assignment. A Score variant of the identical control built that way would separate "reads the position" from "reads the level label", on both layouts. Happy to send it as a small PR if you'd like that split measured.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions