Skip to content

DataSamplingBlock weighted sampling returns duplicate items (random.choices samples with replacement) #15280

Description

@capy-ai

Summary

DataSamplingBlock with sampling_method="weighted" uses random.choices(...), which samples with replacement, so the same item can appear several times in sampled_data and sample_indices. Every other method returns distinct items, and the block says only that it takes sample_size samples from the dataset.

Affected file and line

autogpt_platform/backend/backend/blocks/sampling.py: indices = random.choices(range(data_size), weights=weights, k=input_data.sample_size) (line ~220).

Trigger and steps to reproduce

Run the block with sampling_method="weighted", weight_key="i", sample_size=7 over 12 items such as {"i": n, "g": ...} for n = 0..11. One run gave sample_indices = [9, 9, 8, 10, 7, 10, 6], with indices 9 and 10 repeated.

Expected vs actual

  • Expected: sample_size distinct items, as for the other methods (or an explicit "with replacement" option and description).
  • Actual: repeated items. The returned list can contain fewer unique items than sample_size.

Severity

Low. The result is wrong only relative to what the block's description and the other methods imply. It is not documented either way, so a maintainer may want to treat this as a documentation fix.

How I confirmed it

I ran DataSamplingBlock().run() directly in a randomized loop and saw duplicates in the weighted output. A zero total weight also raises a raw ValueError('Total of weights must be greater than zero') from the stdlib; I did not file that separately.

Fix feasibility

Yes, but it needs a decision: switch to a weighted sample without replacement, or document the with-replacement behaviour.

Activity

  1. ntindle commented on Oct 8, 2026

    @ntindle
    Member

    Backfill validation: reproduced

    Image: significantgravitas/autogpt:latest @ sha256:122929723f57b8927af016d606ae21b21c03a012c2de1d9c30b1c8359e9157f1 (Hub v0.8.3 / sha-73cae306b4f6b197d2e1eaaa3162c326ec0ab076 / oci_revision 73cae306b4f6b197d2e1eaaa3162c326ec0ab076), pulled 2026-10-08 13:08 CT (ok_up_to_date), run sweep-20261008T1808.

    Verdict: reproduced (ran in-image backend code)

    What we checked

    sampling_method="weighted", weight_key="i", sample_size=7 over 12 items returned duplicate indices in 48 of 50 runs (e.g. [10, 11, 6, 5, 6, 10, 9]).

    Limits

    Whether this is a bug or a documentation gap is a maintainer call.

    Suggestion

    Use weighted sampling without replacement, or document that it samples with replacement.

    Host evidence: /workspace/autogpt-backfill/evidence/15280/sweep-20261008T1808/ (host-local)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions