Summary
DataSamplingBlock with sampling_method="weighted" uses random.choices(...), which samples with replacement, so the same item can appear several times in sampled_data and sample_indices. Every other method returns distinct items, and the block says only that it takes sample_size samples from the dataset.
Affected file and line
autogpt_platform/backend/backend/blocks/sampling.py: indices = random.choices(range(data_size), weights=weights, k=input_data.sample_size) (line ~220).
Trigger and steps to reproduce
Run the block with sampling_method="weighted", weight_key="i", sample_size=7 over 12 items such as {"i": n, "g": ...} for n = 0..11. One run gave sample_indices = [9, 9, 8, 10, 7, 10, 6], with indices 9 and 10 repeated.
Expected vs actual
- Expected:
sample_size distinct items, as for the other methods (or an explicit "with replacement" option and description).
- Actual: repeated items. The returned list can contain fewer unique items than
sample_size.
Severity
Low. The result is wrong only relative to what the block's description and the other methods imply. It is not documented either way, so a maintainer may want to treat this as a documentation fix.
How I confirmed it
I ran DataSamplingBlock().run() directly in a randomized loop and saw duplicates in the weighted output. A zero total weight also raises a raw ValueError('Total of weights must be greater than zero') from the stdlib; I did not file that separately.
Fix feasibility
Yes, but it needs a decision: switch to a weighted sample without replacement, or document the with-replacement behaviour.
Summary
DataSamplingBlockwithsampling_method="weighted"usesrandom.choices(...), which samples with replacement, so the same item can appear several times insampled_dataandsample_indices. Every other method returns distinct items, and the block says only that it takessample_sizesamples from the dataset.Affected file and line
autogpt_platform/backend/backend/blocks/sampling.py:indices = random.choices(range(data_size), weights=weights, k=input_data.sample_size)(line ~220).Trigger and steps to reproduce
Run the block with
sampling_method="weighted",weight_key="i",sample_size=7over 12 items such as{"i": n, "g": ...}for n = 0..11. One run gavesample_indices = [9, 9, 8, 10, 7, 10, 6], with indices 9 and 10 repeated.Expected vs actual
sample_sizedistinct items, as for the other methods (or an explicit "with replacement" option and description).sample_size.Severity
Low. The result is wrong only relative to what the block's description and the other methods imply. It is not documented either way, so a maintainer may want to treat this as a documentation fix.
How I confirmed it
I ran
DataSamplingBlock().run()directly in a randomized loop and saw duplicates in the weighted output. A zero total weight also raises a rawValueError('Total of weights must be greater than zero')from the stdlib; I did not file that separately.Fix feasibility
Yes, but it needs a decision: switch to a weighted sample without replacement, or document the with-replacement behaviour.