Getting Started with DPrivStats

library(DPrivStats)
set.seed(42)

This vignette walks through a basic differentially private (DP) analysis workflow: loading sensitive microdata, managing a privacy budget, releasing DP descriptive statistics, and running a DP hypothesis test.

Data

We use a simulated census-like microdata set bundled with the package:

data(example_microdata)
head(example_microdata)
#>   education age    hours region income
#> 1        16  31 34.22962   East  27369
#> 2        11  56 40.69230  North  26722
#> 3         7  49 35.77607  North  13144
#> 4        15  29 33.49301  North  42188
#> 5         7  56 36.17466  South  30107
#> 6        13  43 38.61261   East  20517
nrow(example_microdata)
#> [1] 2000

Privacy budget

All releases draw on a shared privacy budget:

budget <- new_privacy_budget(epsilon = 3.0, delta = 1e-6, composition = "rdp")
budget
#> 
#> Privacy Budget (composition: rdp )
#>   total epsilon:  3.000
#>   spent epsilon:  0.000
#>   accumulated rho: 0.0000
#>   remaining:      3.000

DP descriptive statistics

income_bounds <- c(0, 500000)
m <- dp_mean(example_microdata$income, epsilon = 0.5,
             bounds = income_bounds, mechanism = "laplace")
m
#> 
#> DP mean (laplace mechanism)
#> epsilon = 0.500, delta = 0
#> estimate:
#> [1] 34272.56

The estimate carries its privacy cost and sensitivity with it:

m$sensitivity
#> $l1
#> [1] 250
#> 
#> $l2
#> [1] 250

A DP median via the exponential mechanism, and a DP histogram:

dp_median(example_microdata$income, epsilon = 0.5, bounds = income_bounds,
          n_bins = 200)$estimate
#> [1] 32663.32

hist_fit <- dp_histogram(example_microdata$income, epsilon = 1.0,
                         breaks = seq(0, 500000, by = 100000),
                         normalize = TRUE)
round(hist_fit$estimate, 4)
#>     [0,1e+05] (1e+05,2e+05] (2e+05,3e+05] (3e+05,4e+05] (4e+05,5e+05] 
#>        0.9990        0.0005        0.0002        0.0000        0.0003

DP t-test

Do the two regions differ in income?

sub <- subset(example_microdata, region %in% c("North", "South"))
tt <- dp_t_test(x = sub$income[sub$region == "North"],
                y = sub$income[sub$region == "South"],
                epsilon = 1.0, bounds = income_bounds)
tt
#> 
#>  Differentially Private Two Sample t-test
#> 
#> t = 0.4518, p-value 0.6515
#> df = 1002.2 
#> alternative hypothesis: two.sided 
#> effect size = 1695.8194 (SE = 3753.6521)
#> 
#> privacy cost: epsilon = 1.000, delta = 0

Budget accounting

budget <- spend(budget, 0.5, description = "DP mean income")
budget <- spend(budget, 0.5, description = "DP median income")
budget <- spend(budget, 1.0, description = "DP histogram")
#> Warning: Privacy budget exceeded!
budget <- spend(budget, 1.0, description = "DP t-test")
#> Warning: Privacy budget exceeded!
budget
#> 
#> Privacy Budget (composition: rdp )
#>   total epsilon:  3.000
#>   spent epsilon:  7.127
#>   accumulated rho: 1.2500
#>   remaining:      -4.127