Check your delivery

What you were given is not always what DST documents

Published

October 9, 2026

DST’s variable lists describe a register. Your project gets a delivery: a copy of the registers that someone ordered, converted and placed on your project’s drive. The two can differ. Columns can be missing, keys renamed or empty, years cut off, and registers left out of the parquet folder. None of this causes an error. Your code runs, and the result is quietly wrong.

So before you build a cohort, spend an hour running the checks on this page against your own delivery. Each check comes with an example of what it catches.

The examples are examples. Each one is a problem that turned up in one real project’s delivery. They show what can happen, not what is in yours. Another project, or a newer delivery for the same project, may look completely different. That is the reason to run the checks yourself rather than rely on anyone’s list of known issues, this one included.

All the code reads registers with read_register() from fastreg. Without fastreg, use open_dataset("path/to/<register>/") instead; the rest is the same.

library(fastreg) # read_register(): opens a parquet register by name
library(dplyr) # summarise, count, filter, collect and the pipe

1. Does each register reach the end of your follow-up?

Check the first and last year or date of every register you use against your study period. The code is in Inspect your data.

Example: in one delivery, the LPR2 contact register lpr_adm stopped on 31 December 2018, a full quarter before DST’s own LPR2 ends (31 March 2019). Contacts from those three months existed only in the LPR3 tables, so the standard LPR3 filter quietly removed them. See Extract from LPR for how to check where your LPR2 stops.

2. Does each register have the columns DST documents?

Compare the column names in your delivery with DST’s variable list for that register (Overview of registers links to each one).

read_register("lpr_sksube") %>% # replace with the register you need
  rename_with(tolower) %>%
  colnames() # the columns you actually have

If a column you need is not there, no amount of filtering will bring it back. Ask your data manager whether it can be added to the delivery.

Example: in one delivery, lpr_sksube (examinations and treatments) had only three columns: recnum, d_odto and year. DST documents it with the procedure code c_opr and three more. The table could say that a procedure happened and when, but not which procedure.

3. Are the join keys named as documented?

Joins fail, or join nothing, when a key has a different name than the code expects. Look for the key columns before you write a join.

read_register("t_psyk_adm") %>%
  rename_with(tolower) %>%
  colnames() # is the person key pnr? is the contact key recnum?

Example: in one delivery, the psychiatric LPR2 tables used v_cpr instead of pnr and k_recnum / v_recnum instead of recnum. See Overview of registers for how to rename them only when your delivery needs it.

4. Are the join keys actually filled in?

A key column can exist and still be empty. Count the missing values in every key you plan to join on.

read_register("lpr_a_kontakt") %>% # replace with your table
  rename_with(tolower) %>%
  summarise(
    rows = n(),
    missing_key = sum(is.na(dw_ek_kontakt)) # replace with your join key
  ) %>%
  collect()

Some missing keys can be by design: in LPR3, a procedure attaches either to a contact or to a whole course, and the other key is left empty. A key that is missing on every row is not by design.

Example: in one delivery, a project-specific LPR3 procedure table had the contact key dw_ek_kontakt empty on every single row, so joining it to the contacts returned nothing. The general LPR3 procedure register, lpr_a_procregistrering, had the key filled in and was the way out.

5. Is there one row per key, as documented?

Registers that are documented as one row per person per year (or per family per year) do not always arrive that way. Duplicates make every join grow without an error.

read_register("faik") %>%
  rename_with(tolower) %>%
  count(familie_id, year) %>% # replace with the key DST documents
  filter(n > 1) %>%
  count() %>% # how many keys have more than one row - expect 0
  collect()

Example: in one delivery, FAIK (family income) from 2022 onward also carried pnr and repeated each family’s row once per family member. See Socioeconomic variables for the full check and the fix. BEF can do the same with quarterly snapshots.

6. Is every register in the parquet folder?

A register you applied for is not necessarily converted. Some may only exist as raw SAS files, and the converted one with a similar name may be a different register.

list.files("path/to/parquet-registers/") # replace with your project's path

If a register is missing, ask your data manager where it is. Reading a raw SAS file and converting it are covered in Parquet and fastreg.

Example: in one delivery, the death register dod was only a raw SAS file, while dodsaars was converted and easy to load. dodsaars stops in 2001 (see pitfall 1), so anyone who took the convenient one treated everyone who died after 2001 as alive.

When you find something

Write it down where your project keeps its notes, with the date and the check that found it, and tell your data manager. The next person on the project will otherwise find it again the hard way. Re-run the checks when a new delivery arrives: problems can be fixed, and new ones can appear.

Next steps

Back to top