First extraction
From register to analysis-ready data - step by step
This page shows the complete process from a raw register to a saved dataset ready for analysis. You will see the pattern that recurs in almost every register extraction.
These examples cannot be run on the DST server. The synthetic data package (fiktive) is not available there. You need RStudio installed locally on your computer:
- Download R: cran.r-project.org
- Download RStudio: posit.co/download/rstudio-desktop
- Open RStudio, create a new script (File → New File → R Script), and copy the code there.
When you are ready to work with real register data, you use the same pattern - but on the DST server and with your project’s register data.
The data is synthetic, from fiktive (MIT licence, Sara Schwartz): fictitious persons with the structure and column names of DST registers. load_registers_schema(), generate_background_population(), and generate_registers() return tibbles. You save those with write_parquet(), and you practise open_dataset() on the files. (These local examples use open_dataset(); fastreg’s read_register() is for a configured DST project, not ad-hoc local folders.)
The example extracts from lpr_adm - hospital contacts, dates only, with no diagnosis join. Pulling actual diagnoses brings in rules this page does not cover: the LPR2/LPR3 split, the D-prefix, and choosing diagnosis types. Those are in Understand LPR.
Preparation: install fiktive and generate synthetic data
fiktive is not on CRAN - install directly from GitHub:
install.packages("remotes") # only the first time
remotes::install_github("sara-schwartz/fiktive")library(fiktive) # synthetic DST register data
library(dplyr) # filter, select, mutate, collect
library(arrow) # open_dataset, write_parquet
# Generate synthetic data and save as parquet (done only once).
# One generate_registers() call: every table shares this population, window, and seed.
# The window includes 2015 (the BEF exercise) and both LPR eras used later
# (LPR2 somatic contacts stop around 2019-03; LPR3 diagnoses start in 2019).
schema <- load_registers_schema()
pop <- generate_background_population(n = 60, seed = 1, schema = schema)
tables <- generate_registers(
registers = c("bef", "lpr_adm"),
population = pop,
schema = schema,
from = as.Date("2015-01-01"),
to = as.Date("2020-12-31"),
seed = 1
)
dir.create("synth_data/bef", recursive = TRUE, showWarnings = FALSE) # create folders
dir.create("synth_data/lpr_adm", recursive = TRUE, showWarnings = FALSE)
write_parquet(tables$bef, "synth_data/bef/bef.parquet") # save as parquet
write_parquet(tables$lpr_adm, "synth_data/lpr_adm/lpr_adm.parquet")BEF is quarterly from 2008, and year is an integer, so 2015 holds several rows per person. With n = 60 that year has 240 rows, enough for the sample below.
The path is relative to your working directory. "synth_data/bef/" is created in the folder R is set to work in - check which one with getwd(). To save elsewhere, use a full path: "C:/Users/yourname/projects/synth_data/bef/".
You can now practise open_dataset() exactly as on the DST server.
Step 1 - Define your study population
Every extraction starts with a list of pnr’s - the people you want data for. In practice this comes from your cohort script. Here we build a small practice cohort directly from the BEF register.
step1_cohort.R
# Open BEF - lazy connection, no data in R yet
bef_data <- open_dataset("synth_data/bef/") %>%
rename_with(tolower)
# 200 rows from one snapshot year, not 200 people:
# quarterly BEF has several rows per person in 2015
cohort_pnrs <- bef_data %>%
filter(year == 2015) %>% # take one snapshot year
select(pnr) %>%
collect() %>% # HERE data is fetched into R
slice_sample(n = 200) %>%
pull(pnr)
length(cohort_pnrs) # 200 rows; distinct pnr's can be fewerslice_sample(n = 200) draws rows. Because BEF is quarterly, the same pnr can appear more than once, so this is not a list of 200 people.
In a real project, cohort_pnrs is a vector you built in a previous script and reload with readRDS("path/to/full_cohort.rds") %>% pull(pnr).
The signs in filter(): ,/& = AND, \| = OR, ! = NOT, plus ==, !=, >, >=, %in%. Examples and the parentheses pitfall are in Combining conditions in filter().
Recode BEF variables - what do koen, civst and reg mean?
The BEF register stores koen, civst and reg as codes - not as text. This recoding translates them into analysis-ready variables:
# Continuing from Step 1 - bef_data is already opened with open_dataset()
bef_clean <- bef_data %>%
filter(year == 2015, alder >= 18) %>%
select(pnr, year, foed_dag, koen, civst, reg) %>%
collect() %>% # fetch into R before mutate
mutate(
foed_dato = as.Date(foed_dag), # date format
# koen is numeric: 1 = male, 2 = female
koen_text = if_else(koen == 1, "Male", "Female"),
# civst: marital status
civil_status = case_when(
civst %in% c("G", "P") ~ "Married/partner",
civst %in% c("F", "O", "E", "L") ~ "Divorced/widowed",
civst == "U" ~ "Single",
TRUE ~ NA_character_
),
# reg may be character "81"-"85"; == 81 still matches
region = case_when(
reg == 81 ~ "Region Nordjylland",
reg == 82 ~ "Region Midtjylland",
reg == 83 ~ "Region Syddanmark",
reg == 84 ~ "Region Hovedstaden",
reg == 85 ~ "Region Sjælland",
TRUE ~ NA_character_
),
# partnered from civst; opr_land is all missing in this synthetic BEF
partnered = civst %in% c("G", "P")
)
head(bef_clean)The code lists above are an excerpt, adapted from Anders Aasted Isaksen’s dplyr practice vignette (MIT licence, Steno Diabetes Center Aarhus). The full vignette covers more variables and patterns.
opr_land is entirely missing in this synthetic BEF, so the country-of-origin flag is replaced with partnered from civst.
Step 2 - Extract data from a register
Now we extract hospital contacts from lpr_adm for our cohort. The pattern is always the same: open → filter → select columns → collect.
step2_extraction.R
# Open lpr_adm - lazy connection
lpr_adm <- open_dataset("synth_data/lpr_adm/") %>%
rename_with(tolower)
# Extract: filter BEFORE collect - otherwise the session will crash
contacts <- lpr_adm %>%
semi_join(tibble(pnr = cohort_pnrs), by = "pnr", copy = TRUE) %>% # only our cohort
select(pnr, recnum, d_inddto) %>% # only the columns we use
collect() # HERE data is moved into R
nrow(contacts) # how many contact rows?
head(contacts) # the first six rowsWhat happened?
open_dataset()opened a lazy connection - no data in R yetfilter()andselect()sent instructions to Arrow/DuckDB - still no data in Rcollect()executed the query and fetched only the necessary rows into R
See Extracting data step by step for a detailed explanation of lazy evaluation.
Test on a small sample first. Before running a heavy extraction on the full cohort, test the code on a few people or rows - this catches errors quickly without waiting. E.g. semi_join(tibble(pnr = head(cohort_pnrs, 10)), by = "pnr"), or collect() %>% head(100) while building the code.
Step 3 - Build analysis variables
Add variables with mutate() after collect() - now you are in R and can use all functions.
contacts <- contacts %>%
mutate(
date = as.Date(d_inddto), # explicit date class
year = as.integer(format(date, "%Y")) # year from contact date
)Step 4 - Save and reload
Save with saveRDS() so the next script can reload it without re-running all the extractions.
step4_save.R
saveRDS(contacts, "path/to/extract_contacts.rds") # save to disk - change path to your own folder
# Reload in the next script:
contacts <- readRDS("path/to/extract_contacts.rds")If you do not write a full path, the file is saved in your working directory. Run getwd() to see which folder that is.
The datasets/ folder is stored locally on the DST server only. Intermediate results are repatriated via output control - see 14 - Export and repatriation.
Inspect the result
Right after an extraction you should check that you got what you expected:
head(contacts) # the first six rows - does it look right?
nrow(contacts) # number of rows - as expected?
length(unique(contacts$pnr)) # how many unique individuals?
colSums(is.na(contacts)) # missing values per column
class(contacts$d_inddto) # is the date column Date? (not character)If your extraction includes exclusion steps - e.g. “remove persons with an early diagnosis” - it is good practice to count N for each step. Replace raw and clean with your own variable names:
# Template - replace raw and clean with your own variable names:
cat("Raw extraction: ", nrow(raw), "\n") # all rows before exclusion
cat("After exclusions: ", nrow(clean), "\n") # after each step
cat("Excluded in total: ", nrow(raw) - nrow(clean), "\n") # differenceThese lines cannot be run with the synthetic practice data - they are a template for use when working with your own data and an exclusion sequence. The pattern is repeated for each exclusion step and forms the basis for a STROBE flow diagram (a standardised diagram showing how many were excluded at each step and why; STROBE is the reporting standard for observational studies, CONSORT is for randomised trials). The same attrition counting is used when you build the study population in Phase 10.
This is a quick sanity check. The full toolkit for exploring data - summary(), table(), cross-tables and NA handling - is covered in Phase 7 - Inspect your data.
Next steps
You have now made a complete extraction and saved it. Next steps are to learn to explore data thoroughly:
- Phase 7 - Inspect your data: the full toolkit for understanding what you have
- Extract from LPR: the complete pattern for diagnosis extractions
- Phase 12 - Assemble and prepare the dataset: combine two extracts
Source and adaptation
Step 1 (cohort construction from BEF) is adapted from Anders Aasted Isaksen’s dplyr practice vignette (MIT licence, Steno Diabetes Center Aarhus). The synthetic tables come from fiktive. Steps 2–4 and the checklist are written for this guide.