Summary
The stratified sampling method in DataSamplingBlock raises ValueError: Sample larger than population or is negative for valid inputs, because its top-up step adds samples to a stratum without checking how many items the stratum has.
Affected file and line
autogpt_platform/backend/backend/blocks/sampling.py: the stratum size adjustment loop (lines ~171-180, stratum_sizes[max(stratum_sizes, key=...)] += 1) followed by random.sample(strata[stratum], size) (line ~183).
Trigger and steps to reproduce
Run the block with sampling_method="stratified", stratify_key="g", sample_size=4 and data=[{"g":0},{"g":1},{"g":1},{"g":2},{"g":2}].
Initial sizes are 0 -> max(1, int(0.8)) = 1, 1 -> int(1.6) = 1, 2 -> 1, so the total is 3 and one more sample is needed. The loop increments the first stratum with the largest size (all tie at 1, so stratum 0), which has only 1 item, so random.sample([..1 item..], 2) raises.
Expected vs actual
- Expected: 4 distinct items, with the extra sample going to a stratum that still has capacity (1 + 1 + 2, for example).
- Actual:
ValueError: Sample larger than population or is negative.
Severity
Low to medium. It fails loudly (no wrong data), but valid input crashes with a misleading error. A search over small datasets (n <= 5, up to 3 groups) finds failing cases quickly, so it is not an exotic input.
How I confirmed it
I ran DataSamplingBlock().run() directly over every small dataset up to 5 items with 3 groups; the first failing case is the one above. I did not look at the decrement branch in detail.
Fix feasibility
Yes: when topping up, pick a stratum where size < len(strata[k]), for example by choosing the largest remaining capacity.
Summary
The stratified sampling method in
DataSamplingBlockraisesValueError: Sample larger than population or is negativefor valid inputs, because its top-up step adds samples to a stratum without checking how many items the stratum has.Affected file and line
autogpt_platform/backend/backend/blocks/sampling.py: the stratum size adjustment loop (lines ~171-180,stratum_sizes[max(stratum_sizes, key=...)] += 1) followed byrandom.sample(strata[stratum], size)(line ~183).Trigger and steps to reproduce
Run the block with
sampling_method="stratified",stratify_key="g",sample_size=4anddata=[{"g":0},{"g":1},{"g":1},{"g":2},{"g":2}].Initial sizes are
0 -> max(1, int(0.8)) = 1,1 -> int(1.6) = 1,2 -> 1, so the total is 3 and one more sample is needed. The loop increments the first stratum with the largest size (all tie at 1, so stratum0), which has only 1 item, sorandom.sample([..1 item..], 2)raises.Expected vs actual
ValueError: Sample larger than population or is negative.Severity
Low to medium. It fails loudly (no wrong data), but valid input crashes with a misleading error. A search over small datasets (n <= 5, up to 3 groups) finds failing cases quickly, so it is not an exotic input.
How I confirmed it
I ran
DataSamplingBlock().run()directly over every small dataset up to 5 items with 3 groups; the first failing case is the one above. I did not look at the decrement branch in detail.Fix feasibility
Yes: when topping up, pick a stratum where
size < len(strata[k]), for example by choosing the largest remaining capacity.