Tests multiple sample sizes to determine which achieves desired confidence interval coverage properties. Specifically, it calculates what percentage of 95% confidence intervals (1) exclude zero and (2) include the direct estimate. This helps researchers choose a sample size that provides sufficient precision.

sim_power_N(N.sim = 50, prevalence, p, p.prime, gamma, direct, verbose = TRUE)

sim.power.N(...)

Arguments

N.sim

Integer. Number of Monte Carlo simulations per sample size. Default is 50. Larger values provide more stable estimates but increase computation time.

prevalence

Numeric. True prevalence rate of the sensitive attribute (between 0 and 1).

p

Numeric. Probability for the randomization item in sensitive question.

p.prime

Numeric. Probability for the anchor question (non-sensitive).

gamma

Numeric. Proportion of attentive respondents (between 0 and 1).

direct

Numeric. Direct questioning estimate for comparison purposes.

verbose

Logical. If TRUE (the default), display a progress bar while sample sizes are evaluated.

...

Arguments passed to sim_power_N().

Value

A data frame with three columns:

SampleSize

Sample sizes tested: 100, 500, 1000, 1500, 2000, 2500, 3000

CoverageZero

Percentage of 95% CIs that include zero. Lower values indicate better precision (CIs exclude zero).

CoverageDirect

Percentage of 95% CIs that include the direct estimate. Values near 95% suggest good agreement with direct questioning.

Details

This function is useful for planning studies where researchers want to:

  1. Distinguish the estimated prevalence from zero with high confidence

  2. Obtain narrow confidence intervals for precise estimation

  3. Compare crosswise estimates with direct questioning estimates

For each sample size, the function:

  • Simulates N.sim datasets using the crosswise model

  • Computes bias-corrected estimates with bootstrap 95% CIs

  • Calculates what percentage of CIs contain zero

  • Calculates what percentage of CIs contain the direct estimate

A progress bar displays simulation progress.

Note

  • Low CoverageZero values indicate CIs that reliably exclude zero (good for establishing that prevalence is non-zero)

  • CoverageDirect near 95% suggests consistency between crosswise and direct questioning approaches

  • This function uses 500 bootstrap iterations per simulation

References

Atsusaka and Stevenson (2023). Appendix C5: Sample Size Determination and Parameter Selection. doi:10.1017/pan.2021.43 .

See also

sim_power for the underlying simulation function

Examples

# Find sample size needed to reliably exclude zero
if (FALSE) { # \dontrun{
result <- sim_power_N(
  N.sim = 50,
  prevalence = 0.1,
  p = 0.1,
  p.prime = 0.1,
  gamma = 0.8,
  direct = 0.05
)

print(result)

# Visualize results
plot(result$SampleSize, result$CoverageZero,
     type = "b", xlab = "Sample Size",
     ylab = "% of CIs Including Zero",
     main = "Precision vs Sample Size")
} # }