draw() sets the sample size or sampling fraction, the selection method,
and whatever that method needs: a measure of size for PPS, auxiliary
variables for balanced sampling, coordinates for spatial spread. Every
stage in a sampling design must end with draw().
Usage
draw(
.data,
n = NULL,
frac = NULL,
...,
min_n = NULL,
max_n = NULL,
method = "srswor",
mos = NULL,
prn = NULL,
aux = NULL,
spread = NULL,
round = "up",
control = NULL,
certainty_size = NULL,
certainty_prop = NULL,
certainty_overflow = "error",
on_empty = "error"
)Arguments
- .data
A
sampling_designobject (piped fromsampling_design(),stratify_by(), orcluster_by()).- n
Sample size. For random-size methods (
bernoulli,pps_poisson),nis the expected sample size (converted internally tofrac = n / N). Forpps_poissonthe realized expectation can fall belownwhen large units saturate at probability 1.execute()warns with classsamplyr_warning_poisson_shortfallwhen it falls more than 5% short of what the pool could reach. See selection-methods for the fixed versus random distinction, and usecertainty_sizeorcertainty_propto handle dominant units. Can be:A scalar: applies per stratum (if no
alloc) or as total (ifallocspecified)A named vector: stratum-specific sizes (for single stratification variable)
A data frame: stratum-specific sizes with stratification columns +
ncolumn
- frac
Sampling fraction. Can be:
A scalar: same fraction for all strata
A named vector: stratum-specific fractions
A data frame: stratum-specific fractions with stratification columns +
fraccolumn
Only one of
norfracshould be specified. When the rounded stratum sample size (\(N_h \cdot \text{frac}\)) would be zero, it is floored at 1 so every stratum receives at least one unit.- ...
These dots are for future extensions and must be empty.
.data,n, andfracare what a draw is written as and stay positional. Every argument after...is matched exactly and must be named, so a positional fourth argument is refused rather than taken formin_n.- min_n
Minimum sample size per stratum. When an allocation method (e.g., Neyman, proportional) would assign fewer than
min_nunits to a stratum, that stratum is raised tomin_nand the difference is taken from the remaining strata in proportion to the method's own allocation factors. Commonly set to 2 (minimum for variance estimation) or higher for reliable subgroup estimates. Amin_nabove a stratum's population is structurally capped at that population for without-replacement designs, which makes the stratum a census. Only applies when stratification with an allocation method is used. Default isNULL(no minimum).- max_n
Maximum sample size per stratum. When an allocation method would assign more than
max_nunits to a stratum, that stratum is capped and the surplus is redistributed over the remaining strata in proportion to the method's own allocation factors. Useful for capping dominant strata or managing operational constraints.Neither bound introduces the population cap. An allocation is always capped at the stratum population, so
min_nandmax_nnarrow a range that already exists. See the "Population bounds and redistribution" section ofstratify_by(). For with-replacement methods the number of distinct units is not a bound at all, and these arguments are the only limits that apply. Only applies when stratification with an allocation method is used. Default isNULL(no maximum).To see what a bound does to an allocation before drawing anything, pass the design and the frame to
frame_summary():frame_summary(design, frame, detail = "pool")reports the per-stratum target each bound produces, without selecting or consuming any random numbers.- method
Character string naming the selection method,
"srswor"by default. selection-methods lists all sixteen built-ins with what each one requires and cites the paper behind each."cube"uses auxiliary variables (aux) to balance the sample so that Horvitz-Thompson estimates of auxiliary totals match population totals, and supports equal or unequal (mos) inclusion probabilities. When stratified it uses the stratified cube algorithm (Chauvet 2009). At most two stages may use a balanced-family method."balanced"is retained as a compatibility alias for"cube".Methods registered with
sondage::register_method()use a prefix that identifies their sampling family. Registeredtype = "wor"andtype = "wr"methods use"pps_<name>"(for example,"pps_mymethod"). Registeredtype = "balanced"methods use"balanced_<name>". For the latter,mosis optional.supports_aux = TRUEpermits ordinary balancing variables inaux, andsupports_spread = TRUErequires coordinates inspread. These capabilities are declared when the method is registered insondage.Sample weights are
1 / pik, wherepikis the inclusion expectation (target inclusion probabilities, or expected hits fortype = "wr"methods) that samplyr resolves and hands to the registered method. Registered methods state where they sit in the taxonomy withprobabilitiesat registration:"exact"(the design's true first-order inclusion probabilities, or expected hits, equalpik),"approximate"(honored to a documented approximation, as Pareto sampling does), or"unknown"(the default:pikis a selection weight only, so the weights would be systematically biased).draw()refuses"unknown"methods. Declare the method's tier to use it. The classic trap issample(prob = pik): withreplace = TRUEit yields expected hits exactly equal topik(a validtype = "wr"method), but without replacement its inclusion probabilities differ frompik, so atype = "wor"wrapper's tier really is unknown.- mos
Measure of size variable, specified as a bare column name (unquoted). Required for built-in PPS methods and registered
type = "wor"ortype = "wr"methods named with thepps_prefix. Optional forcube,lpm2,scps, and registeredtype = "balanced"methods named with thebalanced_prefix and when omitted, equal inclusion probabilities are used.- prn
Permanent random number variable for sample coordination, specified as a bare column name (unquoted). Must be a numeric column with values in the open interval (0, 1) and no missing values. Supported methods:
"bernoulli","pps_poisson","pps_sps","pps_pareto". When supplied, the sample is deterministic for a given set of PRN values, enabling coordination across survey waves.- aux
Cube balancing declarations for
method = "cube", or ordinary balancing variables for a registered balanced method that declaressupports_aux = TRUE. Bare numeric columns, such asaux = c(income, pop_density), request approximate Horvitz-Thompson total balance. Abound()marker requests adjacent-integer count bounds for every observed category, for exampleaux = c(income, bound(region), bound(urban_rural)). Use separatebound()calls for separate marginal constraints. Withcluster_by(), ordinary auxiliary values are summed to cluster level, while bound variables must be constant within each cluster.- spread
Spatial coordinates for
method = "lpm2","scps", or a registered balanced method that declaressupports_spread = TRUE, specified as bare numeric columns, for examplespread = c(longitude, latitude). Coordinates must be finite and have no missing values. They should be placed on comparable scales before sampling. Methods declaringsupports_spread = TRUErequire this argument. Withcluster_by(), coordinates must be constant within each cluster.- round
Rounding method when converting
fracto sample sizes. One of:"up"(default): Round up (ceiling). Matches SAS SURVEYSELECT default."down": Round down (floor)."nearest": Round to nearest integer (standard rounding).
This parameter only affects designs using
fracto specify the sampling rate. Whennis specified directly, no rounding occurs. After rounding, a minimum of 1 is enforced per stratum or group.- control
<
data-masking> Variables for sorting the frame before selection. Can be:A single variable:
control = regionMultiple variables:
control = c(region, district)With
serp()for serpentine sorting:control = serp(region, district)With
dplyr::desc()for descending:control = c(region, desc(population))Mixed:
control = c(region, serp(district, commune), desc(size))
See the "Control sorting" section for what the ordering buys and how it interacts with
stratify_by().- certainty_size
For PPS without-replacement methods, units with MOS >= this value are selected with certainty (probability = 1). Can be:
A scalar: same threshold for all strata
A data frame: stratum-specific thresholds with stratification columns
certainty_sizecolumn
Certainty units are removed from the frame before probability sampling, and the remaining sample size is reduced accordingly. Mutually exclusive with
certainty_prop. Equivalent to SAS SURVEYSELECTCERTSIZE=option.- certainty_prop
For PPS without-replacement methods, units whose MOS proportion (MOS_i / sum(MOS)) >= this value are selected with certainty. Can be:
A scalar between 0 and 1 (exclusive), same threshold for all strata
A data frame with stratum-specific thresholds with stratification columns
certainty_propcolumn
Uses iterative selection: after removing certainty units, proportions are recomputed and the check is repeated until no new units qualify. Mutually exclusive with
certainty_size.- certainty_overflow
Controls behavior when certainty units exceed the target sample size
n. One of:"error"(default): Stop with an informative error."allow": Return all certainty units with stage weight 1, even if the resulting sample has more thannunits.
- on_empty
Behavior when a random-size method (
bernoulli,pps_poisson, or a custom method registered withfixed_size = FALSE) selects zero units in a stratum or the whole frame. One of:"error"(default): Stop with an informative error. Zero selections usually indicate a design problem (sampling fraction too small or stratum too small) that should be fixed rather than silently papered over."warn": Issue a warning and keep the empty selection."silent": Keep the empty selection without a message.
An empty selection is a valid realization of a random-size design: it contributes zero to Horvitz-Thompson totals, so estimates from repeated executions remain unbiased. (A fallback that draws a substitute unit would need weights from the combined "draw, then fall back" design. Reusing the SRS or Poisson weights biases HT totals upward.) When an empty stage occurs in a multi-stage design, later stages have nothing to select from and the result is an empty sample.
"warn"and"silent"are intended for simulation and replicated runs. Checknrow()before analyzing a single realization.A replicated execution that produced empty replicates cannot be passed to a later
execute()call (a new phase or a stage continuation): this raises an error of classsamplyr_error_empty_phase_replicaterather than silently skipping the empty replicates, which would condition all downstream results on nonempty realizations. Handle empty replicates explicitly, for example by executing each nonempty replicate separately while accounting for the empty ones in the analysis.
Details
Custom allocation with data frames
For stratum-specific sample sizes or rates, pass a data frame to n or frac.
The data frame must contain:
All stratification variable columns (matching those in
stratify_by())An
ncolumn (for sizes) orfraccolumn (for rates)
Certainty selection
In PPS without-replacement sampling, very large units can have theoretical
inclusion probabilities exceeding 1. Certainty selection handles this by
selecting such units with probability 1 before sampling the remainder.
The output includes a .certainty_k column (where k is the stage number)
indicating which units were certainty selections.
Certainty selection is only available for WOR PPS methods (pps_systematic,
pps_brewer, pps_cps, pps_poisson, pps_sps, pps_pareto). With-replacement methods
(pps_multinomial) and PMR methods (pps_chromy) handle large units
natively through their hit mechanism.
When certainty_overflow = "allow", if more units qualify for certainty
selection than the requested n, all certainty units are returned with
probability 1 (stage weight = 1). No probabilistic sampling is performed in
this case. The resulting sample size will be the number of certainty
units, which exceeds n. In multi-stage designs, the final .weight can
still exceed 1 because it compounds all stage weights.
Certainty with pps_poisson and user-supplied frac. For
pps_poisson, the probabilistic remainder reuses the user-supplied
frac against the remaining (non-certainty) units. That is,
\(\pi_i = \text{frac} \cdot \text{mos}_i \cdot N_r / \sum_{r}\text{mos}_r\)
for the \(N_r\) remaining units, so the expected total sample size is
\(n_{\mathrm{cert}} + \text{frac} \cdot N_r\)
rather than \(\text{frac} \cdot N\). If you need the
expected total to track frac * N, pass an expected n instead and
let samplyr derive the remaining fraction as (n - n_cert) / N_r.
Control sorting
Control sorting orders the sampling frame before selection, providing implicit
stratification. This is particularly effective with systematic and sequential
methods (systematic, pps_systematic, pps_chromy), where it ensures the
sample spreads evenly across the sorted variables.
Serpentine vs nested sorting. Nested (the default,
control = c(var1, var2, var3)) sorts ascending by each variable in turn.
Serpentine (control = serp(var1, var2, var3)) alternates direction at each
hierarchy level, which minimizes the "jumps" between adjacent units and makes
nearby observations more similar. For geographic hierarchies this means the
last district of region 1 is adjacent to the last district of region 2.
Combining with explicit stratification. When both stratify_by() and
control are used, sorting is applied within each stratum, giving explicit
stratification for variance control alongside implicit stratification for
sample spread.
See also
sampling_design() for creating designs,
stratify_by() for stratification,
cluster_by() for clustering,
execute() for running designs,
serp() for serpentine sorting
Other design specification:
add_stage(),
cluster_by(),
sampling_design(),
selection-methods,
stratify_by()
Examples
# Simple random sample of 100 EAs
sampling_design() |>
draw(n = 100) |>
execute(bfa_eas, seed = 1)
#> # A tbl_sample: 100 × 18
#> # Sampling: 1 stage | 100/44,570 units
#> # Weights: 445.7 [445.7, 445.7]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 43475 Est Gnagna Piéla Rural 971 112 8.95
#> 2 6592 Sud-Ouest Noumbiel Kpuéré Rural 88 12 7.89
#> 3 11611 Boucle du … Nayala Yaba Rural 111 15 8.97
#> 4 45236 Centre-Est Boulgou Béguédo Urban 939 201 0.32
#> 5 3549 Est Gourma Fada-N… Rural 263 32 3.4
#> 6 39095 Hauts-Bass… Kenedou… Sindo Rural 37 4 6.09
#> 7 21818 Centre-Est Kourite… Goungu… Rural 21 4 1.07
#> 8 15528 Centre Kadiogo Ouagad… Urban 1207 182 0.14
#> 9 14284 Est Gourma Matiak… Rural 178 21 8.86
#> 10 3433 Est Gourma Fada-N… Rural 285 34 8.55
#> # ℹ 90 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# Systematic sample of 10%
sampling_design() |>
draw(frac = 0.10, method = "systematic") |>
execute(bfa_eas, seed = 123)
#> # A tbl_sample: 4457 × 18
#> # Sampling: 1 stage | 4,457/44,570 units
#> # Weights: 10 [10, 10]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 11761 Boucle du … Bale Bagassi Rural 63 8 8.54
#> 2 11771 Boucle du … Bale Bagassi Rural 234 28 5.91
#> 3 11781 Boucle du … Bale Bagassi Rural 86 10 3.68
#> 4 11791 Boucle du … Bale Bagassi Rural 970 117 1.2
#> 5 11801 Boucle du … Bale Bagassi Rural 102 12 9.29
#> 6 11811 Boucle du … Bale Bagassi Rural 135 16 8.25
#> 7 11821 Boucle du … Bale Bagassi Rural 208 25 8.13
#> 8 11831 Boucle du … Bale Bagassi Rural 92 11 8.97
#> 9 11843 Boucle du … Bale Bagassi Rural 132 16 10.7
#> 10 11853 Boucle du … Bale Bagassi Rural 63 8 6.11
#> # ℹ 4,447 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# PPS sample of EAs using household count
sampling_design() |>
cluster_by(ea_id) |>
draw(n = 50, method = "pps_brewer", mos = households) |>
execute(bfa_eas, seed = 42)
#> # A tbl_sample: 50 × 19
#> # Weights: 970.83 [117.62, 7498.52]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 1192 Boucle du … Bale Boromo Rural 2570 374 2.05
#> 2 36794 Boucle du … Bale Fara Rural 2115 309 1.07
#> 3 41812 Boucle du … Kossi Dokui Rural 718 99 1.24
#> 4 11459 Boucle du … Sourou Toéni Rural 461 68 0.36
#> 5 27089 Cascades Leraba Loumana Rural 491 66 9.54
#> 6 15148 Centre Kadiogo Ouagad… Urban 907 137 0.16
#> 7 15241 Centre Kadiogo Ouagad… Urban 983 148 0.14
#> 8 16887 Centre Kadiogo Ouagad… Urban 914 138 0.1
#> 9 18171 Centre Kadiogo Ouagad… Urban 1349 203 0.39
#> 10 831 Centre-Est Boulgou Bissiga Rural 368 81 3.22
#> # ℹ 40 more rows
#> # ℹ 11 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>, .certainty_1 <lgl>
# Bernoulli sampling with frac (random sample size, expected ~5%)
sampling_design() |>
draw(frac = 0.05, method = "bernoulli") |>
execute(ken_enterprises, seed = 12345)
#> # A tbl_sample: 827 × 14
#> # Sampling: 1 stage | 827/17,004 units
#> # Weights: 20 [20, 20]
#> enterprise_id county region sector size_class employees revenue_millions
#> * <chr> <fct> <fct> <fct> <fct> <int> <dbl>
#> 1 KEN_00010 Nyandarua Central Chemic… Small 10 40.5
#> 2 KEN_00020 Kiambu Central Constr… Small 9 17.4
#> 3 KEN_00023 Kiambu Central Constr… Small 17 17.3
#> 4 KEN_00051 Kirinyaga Central Constr… Small 10 25.5
#> 5 KEN_00063 Murang'a Central Constr… Medium 28 81.9
#> 6 KEN_00065 Murang'a Central Constr… Medium 31 30.5
#> 7 KEN_00077 Murang'a Central Constr… Small 10 14.7
#> 8 KEN_00103 Nyandarua Central Constr… Small 14 18.4
#> 9 KEN_00117 Nyandarua Central Constr… Small 12 31.1
#> 10 KEN_00146 Nyeri Central Constr… Small 9 8.3
#> # ℹ 817 more rows
#> # ℹ 7 more variables: year_established <int>, exporter <lgl>, .weight <dbl>,
#> # .sample_id <int>, .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# Bernoulli sampling with expected n (converted to frac = 500/N)
sampling_design() |>
draw(n = 500, method = "bernoulli") |>
execute(bfa_eas, seed = 42)
#> # A tbl_sample: 526 × 18
#> # Sampling: 1 stage | 526/44,570 units
#> # Weights: 89.14 [89.14, 89.14]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 11781 Boucle du … Bale Bagassi Rural 86 10 3.68
#> 2 36665 Boucle du … Bale Fara Rural 518 76 8.89
#> 3 36750 Boucle du … Bale Fara Rural 107 16 4.14
#> 4 9032 Boucle du … Bale Poura Urban 432 82 4.38
#> 5 9047 Boucle du … Bale Poura Urban 181 34 4.7
#> 6 9053 Boucle du … Bale Poura Urban 436 83 3.19
#> 7 39999 Boucle du … Banwa Balavé Rural 1820 287 2.04
#> 8 34012 Boucle du … Banwa Sami Rural 135 20 5.63
#> 9 9680 Boucle du … Banwa Sanaba Rural 59 8 8.98
#> 10 23812 Boucle du … Banwa Solenzo Rural 679 91 8.85
#> # ℹ 516 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# Stratified with different sizes per stratum (data frame)
region_sizes <- data.frame(
region = levels(bfa_eas$region),
n = c(20, 12, 25, 18, 22, 16, 14, 15, 20, 18, 12, 10, 8)
)
sampling_design() |>
stratify_by(region) |>
draw(n = region_sizes) |>
execute(bfa_eas, seed = 123)
#> # A tbl_sample: 210 × 18
#> # Sampling: 1 stage | 210/44,570 units
#> # Weights: 212.24 [115.14, 414.4]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 33181 Boucle du … Kossi Nouna Rural 133 17 0.07
#> 2 33229 Boucle du … Kossi Nouna Rural 48 6 4.53
#> 3 21506 Boucle du … Kossi Doumba… Rural 485 83 4.82
#> 4 45264 Boucle du … Bale Pâ Rural 825 95 0.66
#> 5 26201 Boucle du … Sourou Di Rural 95 17 21.1
#> 6 26077 Boucle du … Mouhoun Dédoug… Rural 176 25 8.96
#> 7 25510 Boucle du … Kossi Bombor… Rural 246 36 0.44
#> 8 23774 Boucle du … Banwa Solenzo Rural 158 21 2.9
#> 9 8332 Boucle du … Mouhoun Ouarko… Rural 55 8 5.68
#> 10 43713 Boucle du … Mouhoun Safané Rural 97 14 6.46
#> # ℹ 200 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <dbl>
# Stratified with different rates per stratum (named vector)
sampling_design() |>
stratify_by(size_class) |>
draw(frac = c(Small = 0.02, Medium = 0.10, Large = 0.50)) |>
execute(ken_enterprises, seed = 42)
#> # A tbl_sample: 1330 × 14
#> # Sampling: 1 stage | 1,330/17,004 units
#> # Weights: 12.78 [2, 50]
#> enterprise_id county region sector size_class employees revenue_millions
#> * <chr> <fct> <fct> <fct> <fct> <int> <dbl>
#> 1 KEN_08403 Nairobi Nairobi Other… Medium 64 94.6
#> 2 KEN_14993 Siaya Nyanza… Other… Medium 73 141.
#> 3 KEN_08163 Nairobi Nairobi Other… Medium 53 68.2
#> 4 KEN_03931 Isiolo East a… Other… Medium 95 158.
#> 5 KEN_04478 Nairobi Nairobi Const… Medium 60 52.1
#> 6 KEN_02158 Tana River Coast Other… Medium 43 33.9
#> 7 KEN_06792 Nairobi Nairobi Hotel… Medium 41 81.7
#> 8 KEN_14492 Migori Nyanza… Const… Medium 54 57.8
#> 9 KEN_01181 Kilifi Coast Const… Medium 71 213.
#> 10 KEN_15690 Laikipia Rift V… Hotel… Medium 20 39.8
#> # ℹ 1,320 more rows
#> # ℹ 7 more variables: year_established <int>, exporter <lgl>, .weight <dbl>,
#> # .sample_id <int>, .stage <int>, .weight_1 <dbl>, .fpc_1 <dbl>
# Neyman allocation with minimum 2 per stratum (for variance estimation)
sampling_design() |>
stratify_by(region, alloc = "neyman", variance = bfa_eas_variance) |>
draw(n = 150, min_n = 2) |>
execute(bfa_eas, seed = 2026)
#> # A tbl_sample: 150 × 18
#> # Sampling: 1 stage | 150/44,570 units
#> # Weights: 297.13 [278.67, 332.4]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 34795 Boucle du … Sourou Tougan Rural 45 7 14.0
#> 2 44496 Boucle du … Mouhoun Tchéri… Rural 180 29 16.6
#> 3 9624 Boucle du … Banwa Sanaba Rural 177 23 6.56
#> 4 7012 Boucle du … Kossi Madouba Rural 387 48 6.18
#> 5 44420 Boucle du … Mouhoun Tchéri… Rural 63 10 22.3
#> 6 515 Boucle du … Kossi Barani Rural 651 85 1.11
#> 7 1144 Boucle du … Bale Boromo Rural 75 11 8.92
#> 8 8378 Boucle du … Bale Ouri Rural 693 98 9.26
#> 9 4909 Boucle du … Sourou Kassoum Rural 662 90 0.65
#> 10 21094 Boucle du … Kossi Bouras… Rural 25 4 5.1
#> # ℹ 140 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <dbl>
# Proportional allocation with min and max bounds
sampling_design() |>
stratify_by(region, alloc = "proportional") |>
draw(n = 200, min_n = 10, max_n = 50) |>
execute(bfa_eas, seed = 1)
#> # A tbl_sample: 200 × 18
#> # Sampling: 1 stage | 200/44,570 units
#> # Weights: 222.85 [161.2, 240.7]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 9648 Boucle du … Banwa Sanaba Rural 279 35 8.37
#> 2 11547 Boucle du … Sourou Toéni Rural 49 7 21.4
#> 3 41824 Boucle du … Kossi Dokui Rural 75 10 15.8
#> 4 11012 Boucle du … Banwa Tansila Rural 592 71 0.95
#> 5 32308 Boucle du … Sourou Lanfiè… Rural 57 8 9.49
#> 6 7017 Boucle du … Kossi Madouba Rural 599 74 0.74
#> 7 36700 Boucle du … Bale Fara Rural 402 59 6.36
#> 8 11611 Boucle du … Nayala Yaba Rural 111 15 8.97
#> 9 8342 Boucle du … Mouhoun Ouarko… Rural 94 13 8.1
#> 10 11626 Boucle du … Nayala Yaba Rural 56 7 8.31
#> # ℹ 190 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <dbl>
# Control sorting with serpentine ordering (implicit stratification)
sampling_design() |>
draw(n = 100, method = "systematic",
control = serp(region, province)) |>
execute(bfa_eas, seed = 2)
#> # A tbl_sample: 100 × 18
#> # Sampling: 1 stage | 100/44,570 units
#> # Weights: 445.7 [445.7, 445.7]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 11843 Boucle du … Bale Bagassi Rural 132 16 10.7
#> 2 8876 Boucle du … Bale Pompoï Rural 55 8 8.62
#> 3 34070 Boucle du … Banwa Sami Rural 24 4 8.12
#> 4 24054 Boucle du … Banwa Solenzo Rural 123 17 8.82
#> 5 21070 Boucle du … Kossi Bouras… Rural 115 18 9.87
#> 6 37405 Boucle du … Kossi Kombori Rural 308 39 9.76
#> 7 21009 Boucle du … Mouhoun Bondok… Rural 556 73 9.16
#> 8 5672 Boucle du … Mouhoun Kona Rural 222 35 8.98
#> 9 44440 Boucle du … Mouhoun Tchéri… Rural 55 9 5.53
#> 10 11656 Boucle du … Nayala Yaba Rural 1292 170 1.35
#> # ℹ 90 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# Control sorting with nested (standard) ordering
sampling_design() |>
draw(n = 100, method = "systematic",
control = c(region, province)) |>
execute(bfa_eas, seed = 3)
#> # A tbl_sample: 100 × 18
#> # Sampling: 1 stage | 100/44,570 units
#> # Weights: 445.7 [445.7, 445.7]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 11833 Boucle du … Bale Bagassi Rural 35 4 7.15
#> 2 8620 Boucle du … Bale Pâ Rural 52 6 1.46
#> 3 34062 Boucle du … Banwa Sami Rural 286 43 7.68
#> 4 24046 Boucle du … Banwa Solenzo Rural 24 3 6.37
#> 5 25526 Boucle du … Kossi Bombor… Rural 152 22 0.06
#> 6 37398 Boucle du … Kossi Kombori Rural 201 25 5.89
#> 7 21002 Boucle du … Mouhoun Bondok… Rural 791 104 9.13
#> 8 5663 Boucle du … Mouhoun Kona Rural 180 28 7.94
#> 9 44432 Boucle du … Mouhoun Tchéri… Rural 502 80 23.6
#> 10 11649 Boucle du … Nayala Yaba Rural 526 69 7.66
#> # ℹ 90 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# Combined explicit stratification with control sorting within strata
sampling_design() |>
stratify_by(urban_rural) |>
draw(n = 50, method = "systematic",
control = serp(region, province)) |>
execute(bfa_eas, seed = 25)
#> # A tbl_sample: 100 × 18
#> # Sampling: 1 stage | 100/44,570 units
#> # Weights: 445.7 [137.66, 753.74]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 36744 Boucle du … Bale Fara Rural 565 83 9.07
#> 2 9728 Boucle du … Banwa Sanaba Rural 69 9 6.85
#> 3 25519 Boucle du … Kossi Bombor… Rural 240 35 4.62
#> 4 10541 Boucle du … Kossi Sônô Rural 738 110 0.6
#> 5 8318 Boucle du … Mouhoun Ouarko… Rural 160 22 7.76
#> 6 11674 Boucle du … Nayala Yaba Rural 1353 178 1.3
#> 7 34832 Boucle du … Sourou Tougan Rural 33 5 8.81
#> 8 20641 Cascades Comoe Banfora Rural 996 114 0.24
#> 9 15093 Cascades Comoe Niango… Rural 126 15 8.87
#> 10 10357 Cascades Comoe Sidéra… Rural 31 5 8.64
#> # ℹ 90 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>
# PPS with certainty selection (absolute threshold)
# Large EAs selected with certainty, rest sampled with PPS
sampling_design() |>
stratify_by(region) |>
draw(n = 100, method = "pps_brewer", mos = households,
certainty_size = 800) |>
execute(bfa_eas, seed = 3)
#> # A tbl_sample: 1300 × 19
#> # Sampling: 1 stage | 1,300/44,570 units
#> # Weights: 35.46 [1, 1097.49]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 39984 Boucle du … Banwa Balavé Rural 193 30 0.3
#> 2 11674 Boucle du … Nayala Yaba Rural 1353 178 1.3
#> 3 12747 Boucle du … Kossi Djibas… Rural 679 97 0.37
#> 4 11053 Boucle du … Banwa Tansila Rural 651 78 2.31
#> 5 26118 Boucle du … Mouhoun Dédoug… Rural 683 97 0.54
#> 6 26130 Boucle du … Mouhoun Dédoug… Rural 415 59 0.13
#> 7 8916 Boucle du … Bale Pompoï Rural 916 135 0.73
#> 8 23995 Boucle du … Banwa Solenzo Rural 887 119 0.59
#> 9 26018 Boucle du … Mouhoun Dédoug… Rural 985 140 1
#> 10 5658 Boucle du … Mouhoun Kona Rural 616 97 8.53
#> # ℹ 1,290 more rows
#> # ℹ 11 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>, .certainty_1 <lgl>
# PPS with certainty selection (proportional threshold)
# EAs with >= 10% of stratum total selected with certainty
sampling_design() |>
stratify_by(region) |>
draw(n = 100, method = "pps_systematic", mos = households,
certainty_prop = 0.10) |>
execute(bfa_eas, seed = 321)
#> # A tbl_sample: 1300 × 19
#> # Sampling: 1 stage | 1,300/44,570 units
#> # Weights: 33.25 [2.16, 661.2]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 11818 Boucle du … Bale Bagassi Rural 37 4 6.38
#> 2 29503 Boucle du … Bale Bana Rural 530 69 0.45
#> 3 1136 Boucle du … Bale Boromo Rural 3613 526 4.23
#> 4 1177 Boucle du … Bale Boromo Rural 1383 201 2.06
#> 5 1207 Boucle du … Bale Boromo Rural 1465 213 1.67
#> 6 36701 Boucle du … Bale Fara Rural 638 93 0.81
#> 7 36734 Boucle du … Bale Fara Rural 1114 163 1.07
#> 8 36769 Boucle du … Bale Fara Rural 995 145 1.41
#> 9 8390 Boucle du … Bale Ouri Rural 682 97 8.14
#> 10 8435 Boucle du … Bale Ouri Rural 624 88 0.61
#> # ℹ 1,290 more rows
#> # ℹ 11 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>, .certainty_1 <lgl>
# Stratum-specific certainty thresholds (data frame)
cert_thresholds <- data.frame(
region = levels(bfa_eas$region),
certainty_size = c(700, 450, 800, 850, 750, 800, 550,
450, 700, 950, 750, 600, 480)
)
sampling_design() |>
stratify_by(region) |>
draw(n = 100, method = "pps_brewer", mos = households,
certainty_size = cert_thresholds) |>
execute(bfa_eas, seed = 424)
#> # A tbl_sample: 1300 × 19
#> # Sampling: 1 stage | 1,300/44,570 units
#> # Weights: 35.78 [1, 847.49]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 9046 Boucle du … Bale Poura Urban 4132 784 3.25
#> 2 34802 Boucle du … Sourou Tougan Rural 80 12 7.67
#> 3 12839 Boucle du … Kossi Djibas… Rural 1300 185 1.74
#> 4 39958 Boucle du … Banwa Balavé Rural 684 108 0.88
#> 5 9914 Boucle du … Bale Siby Rural 166 25 3.44
#> 6 23777 Boucle du … Banwa Solenzo Rural 485 65 8.06
#> 7 25901 Boucle du … Mouhoun Dédoug… Rural 901 128 7.84
#> 8 34036 Boucle du … Banwa Sami Rural 357 54 6.67
#> 9 44351 Boucle du … Mouhoun Tchéri… Rural 234 37 8.31
#> 10 26001 Boucle du … Mouhoun Dédoug… Rural 183 26 0.06
#> # ℹ 1,290 more rows
#> # ℹ 11 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>, .certainty_1 <lgl>