The basic workflow
samplyr separates a sampling design from the frame on
which it runs. You describe the design first, validate or inspect it,
and then execute it against one or more frames.
The smallest complete example is a simple random sample without replacement.
data(bfa_eas)
design <- sampling_design(title = "Burkina Faso EA sample") |>
draw(n = 100)
sample <- execute(design, bfa_eas, seed = 24)
sample
#> # A tbl_sample: 100 × 18 | Burkina Faso EA sample
#> # Sampling: 1 stage | 100/44,570 units
#> # Weights: 445.7 [445.7, 445.7]
#> ea_id region province commune urban_rural population households area_km2
#> * <int> <fct> <fct> <fct> <fct> <int> <int> <dbl>
#> 1 5866 Centre-Nord Bam Kongou… Urban 978 146 0.25
#> 2 24905 Nord Passore Arbollé Rural 280 32 0.51
#> 3 5145 Centre-Nord Sanmate… Kaya Urban 959 156 0.28
#> 4 32861 Est Gnagna Mani Rural 162 22 0.12
#> 5 4797 Sud-Ouest Poni Kampti Rural 147 23 6.59
#> 6 25986 Boucle du … Mouhoun Dédoug… Rural 849 120 1.49
#> 7 16401 Centre Kadiogo Ouagad… Urban 815 123 0.09
#> 8 13189 Nord Zondoma Goursi Rural 211 24 0.23
#> 9 16345 Centre Kadiogo Ouagad… Urban 928 140 0.21
#> 10 30414 Sahel Seno Dori Rural 264 34 7.02
#> # ℹ 90 more rows
#> # ℹ 10 more variables: pop_density <dbl>, longitude <dbl>, latitude <dbl>,
#> # remoteness <fct>, fieldwork_cost <int>, .weight <dbl>, .sample_id <int>,
#> # .stage <int>, .weight_1 <dbl>, .fpc_1 <int>The frame has one row per enumeration area. The result keeps its frame columns and adds the design information needed for analysis.
library(dplyr)
sample |>
as_tibble() |>
select(ea_id, region, households, .weight, .weight_1) |>
slice_head(n = 6)
#> # A tibble: 6 × 5
#> ea_id region households .weight .weight_1
#> <int> <fct> <int> <dbl> <dbl>
#> 1 5866 Centre-Nord 146 446. 446.
#> 2 24905 Nord 32 446. 446.
#> 3 5145 Centre-Nord 156 446. 446.
#> 4 32861 Est 22 446. 446.
#> 5 4797 Sud-Ouest 23 446. 446.
#> 6 25986 Boucle du Mouhoun 120 446. 446..weight is the overall design weight. In a one-stage
sample it equals .weight_1. A multistage sample also
carries one weight column per stage and multiplies them into
.weight.
The design grammar
Six functions cover the ordinary workflow.
| Function | Role |
|---|---|
sampling_design() |
Start a reusable design |
stratify_by() |
Divide a stage into sampling strata |
cluster_by() |
Name the units selected as clusters |
draw() |
Choose the size and selection method |
add_stage() |
Start another sampling stage |
execute() |
Run the design against its frame or frames |
Column names are stored in the design and resolved when a frame is supplied. This lets the same design be validated and reused against compatible frame vintages.
stratified_design <- sampling_design(title = "Stratified EA sample") |>
stratify_by(region, alloc = "proportional") |>
draw(n = 260)
validate_frame(stratified_design, bfa_eas)validate_frame() returns invisibly when the required
columns and values are usable. It reports problems before selection
begins.
Stratified sampling
Stratification samples independently within groups. Here
n = 260 is a total allocated across regions in proportion
to their frame sizes.
stratified_sample <- execute(stratified_design, bfa_eas, seed = 25)
stratified_sample |>
count(region, name = "n_sampled")
#> # A tibble: 13 × 2
#> region n_sampled
#> <fct> <int>
#> 1 Boucle du Mouhoun 29
#> 2 Cascades 15
#> 3 Centre 23
#> 4 Centre-Est 17
#> 5 Centre-Nord 20
#> 6 Centre-Ouest 22
#> 7 Centre-Sud 9
#> 8 Est 32
#> 9 Hauts-Bassins 28
#> 10 Nord 17
#> 11 Plateau-Central 10
#> 12 Sahel 24
#> 13 Sud-Ouest 14If stratify_by() has no allocation rule, a scalar
n is interpreted per stratum. You can also pass a named
vector or a svyplan::n_alloc() result to
draw() when each stratum needs its own take.
Use equal allocation when comparable precision within each stratum is the goal. Use proportional allocation for a self-weighting element design when the other design features permit it. Neyman and cost-weighted allocations need prior variability or cost information and optimize the criterion supplied for that planning variable. They are not universal optima for every outcome and domain.
Sample-size bounds can protect small reporting domains.
bounded_design <- sampling_design() |>
stratify_by(region, alloc = "proportional") |>
draw(n = 260, min_n = 12, max_n = 30)
bounded_sample <- execute(bounded_design, bfa_eas, seed = 26)
bounded_sample |>
count(region, name = "n_sampled")
#> # A tibble: 13 × 2
#> region n_sampled
#> <fct> <int>
#> 1 Boucle du Mouhoun 29
#> 2 Cascades 14
#> 3 Centre 22
#> 4 Centre-Est 17
#> 5 Centre-Nord 20
#> 6 Centre-Ouest 21
#> 7 Centre-Sud 12
#> 8 Est 30
#> 9 Hauts-Bassins 28
#> 10 Nord 17
#> 11 Plateau-Central 12
#> 12 Sahel 24
#> 13 Sud-Ouest 14For without-replacement sampling, the population of a stratum remains
a hard upper bound. If a requested allocation cannot be reached,
execute() reports the affected pools and uses the
attainable take.
Cluster and PPS sampling
cluster_by() changes the sampling unit. The next example
selects enumeration areas with probability related to their household
counts.
pps_design <- sampling_design(title = "PPS EA sample") |>
cluster_by(ea_id) |>
draw(n = 80, method = "pps_brewer", mos = households)
pps_sample <- execute(pps_design, bfa_eas, seed = 27)
nrow(pps_sample)
#> [1] 80mos is the measure of size. A large measure gives a unit
a larger selection probability. If a resolved probability reaches one,
that unit is selected with certainty and contributes no variance at that
stage. Certainty at one stage does not make the final multistage weight
equal to one.
pps_brewer is one fixed-size PPS method. Method choice
also depends on the second-order information required for variance
estimation, whether a random sample size is acceptable, and whether
permanent random numbers are needed. The comparison is in
?selection-methods.
A two-stage sample
Multistage designs select units within units chosen earlier. This small example selects EAs first and then households from a separate listing frame.
ea_frame <- bfa_eas |>
slice_head(n = 200)
household_frame <- data.frame(
ea_id = rep(ea_frame$ea_id, each = 12),
household_id = sequence(rep(12, nrow(ea_frame))),
outcome = rnorm(12 * nrow(ea_frame))
)
two_stage_design <- sampling_design(title = "EA and household sample") |>
cluster_by(ea_id) |>
draw(n = 20) |>
add_stage() |>
draw(n = 5)
household_sample <- execute(
two_stage_design,
ea_frame,
household_frame,
seed = 28
)
c(
eas = n_distinct(household_sample$ea_id),
households = nrow(household_sample)
)
#> eas households
#> 20 100Frames are mapped to stages by position. samplyr
restricts the household listing to the selected EAs, then compounds the
conditional stage weights.
household_sample |>
select(ea_id, household_id, .weight_1, .weight_2, .weight) |>
slice_head(n = 6)
#> # A tbl_sample: 6 × 5 | EA and household sample
#> # Modified: columns, rows
#> # Weights: 24 [24, 24]
#> ea_id household_id .weight_1 .weight_2 .weight
#> * <int> <int> <dbl> <dbl> <dbl>
#> 1 11783 12 10 2.4 24
#> 2 11783 9 10 2.4 24
#> 3 11783 3 10 2.4 24
#> 4 11783 7 10 2.4 24
#> 5 11783 1 10 2.4 24
#> 6 11785 7 10 2.4 24When later listings become available only after stage 1 fieldwork,
execute the same design in steps. Pass stages = 1 for the
first selection, then continue from that unmodified partial sample.
selected_eas <- execute(two_stage_design, ea_frame, stages = 1, seed = 28)
continued_sample <- execute(selected_eas, household_frame)
identical(household_sample$household_id, continued_sample$household_id)
#> [1] FALSEOne seed around the single-call execution covers the whole random sequence. The continuation stores the required RNG state, so the two routes agree.
Fractions, replacement, and ordering
draw() also accepts frac. Fixed-size
methods convert the fraction to an integer take, while Bernoulli
sampling keeps a random realized size.
fraction_sample <- sampling_design() |>
draw(frac = 0.01) |>
execute(bfa_eas, seed = 29)
nrow(fraction_sample)
#> [1] 446With-replacement methods may select the same population unit several
times. Their samples carry .draw_1, .draw_2,
and later-stage equivalents so every draw occurrence remains
distinct.
Systematic and sequential methods depend on frame order. Use
control when an explicit ordering is part of the design.
Their variance properties depend on that ordering, so an SRS variance
approximation for systematic selection need not be conservative.
Inspecting the result
A tbl_sample is a tibble with a design and an execution
receipt attached. Ordinary one-to-one joins and new analysis columns
preserve it. Removing or duplicating rows changes the realized sample
and prevents survey export.
summary(stratified_sample)
#> ── Sample Summary: Stratified EA sample ────────────────────────────────────────
#>
#> ℹ n = 260 of 44,570 | stages = 1/1 | seed = 25
#>
#> ── Stage 1 ─────────────────────────────────────────────────────────────────────
#> • srswor, by region (proportional)
#> • 13 strata: N_h 1,612-5,505, n_h 9-32, f_h 0.0056-0.0060
#>
#> ── Weights ─────────────────────────────────────────────────────────────────────
#> • Mean 171.42 [166.2, 179.11] | CV 0.01 | Kish DEFF 1 | n_eff 260frame_summary() reports what the execution resolved in
each selection pool.
frame_summary(stratified_sample, detail = "pool") |>
select(stage, region, N, n_target, n_realized) |>
slice_head(n = 6)
#> # A tibble: 6 × 5
#> stage region N n_target n_realized
#> <int> <fct> <dbl> <dbl> <dbl>
#> 1 1 Boucle du Mouhoun 5009 29 29
#> 2 1 Cascades 2508 15 15
#> 3 1 Centre 3888 23 23
#> 4 1 Centre-Est 2941 17 17
#> 5 1 Centre-Nord 3402 20 20
#> 6 1 Centre-Ouest 3723 22 22For domain estimates, convert the complete sample first and subset
the survey design afterwards. See
vignette("survey-analysis") for that workflow.
Where the specialized methods live
The getting-started path does not require every selection family.
| Need | Next reference |
|---|---|
| Choose a sample size or allocation | vignette("survey-planning") |
| Export weights and estimate domains | vignette("survey-analysis") |
| Coordinate repeated samples with PRNs | vignette("sampling-coordination") |
| Build a rotating panel | vignette("rotating-panels") |
| Understand probability and variance contracts | vignette("design-semantics") |
| Save or replay a design | vignette("serialization") |
Balanced sampling, spatially balanced sampling, custom methods, hard
balance constraints, replicate execution, and detailed joint-probability
calculations remain available through ?selection-methods
and the statistical-contract article. They are useful tools, but they
are not prerequisites for drawing a valid first sample.
A compact checklist
- Make sure one row represents the unit selected at the current stage.
- Declare strata, clusters, and selection methods in their sampling order.
- Validate the design against the intended frame vintage.
- Use a recorded seed when the selection must be reproducible.
- Keep the returned design columns unchanged.
- Archive the executed sample and its frame provenance.
The next article for most users is
vignette("survey-planning") when the take is not yet known,
or vignette("survey-analysis") when data collection is
complete.