Summary of the Anonymization Workflow

  • Understand the anonymization workflow.

This overview summarizes the full anonymization workflow. Use it as a reference or checklist when you work through your own data. Keep in mind that not every dataset needs every step in full; how much effort is appropriate depends on your data’s risk (see Does It Always Take This Much Work?).

1
Implement Data Protection Privacy by design before and during collection: minimize data, plan storage/access, write a Data Management Plan.
▼
2
Collect Data & Handle Identifiers Use informed consent and pseudonymize early. Remove direct identifiers from every copy, store the pseudonym key separately, and work from a secure identifier-free copy.
▼
3
Run Analysis Conduct your analysis, writing, and collaborator sharing on the identifier-free working copy.
▼
4
Analyze Attack Scenarios Consider adversaries and disclosure risks (identity, attribute, inference, membership); set a k-anonymity goal.
▼
5
Calculate Disclosure Risk & Utility Measure k-anonymity and risk metrics; establish a utility baseline (e.g., using sdcMicro).
▼
6
Choose Anonymization Measures Select a technique: non-perturbative, perturbative, de-associative, or synthetic. When in doubt, aim for k-anonymity.
▼
7
Apply Anonymization Techniques Apply chosen techniques step by step; track changes with sdcMicro.
▼
8
Recalculate Risk, Utility & Statistics Re-measure after each step. If the balance is not satisfactory, return to Step 4 and iterate.
▼
9
Document the Process Internal (auditing) and external (data dictionary) documentation; review scripts; state in README that results may not be fully reproducible.
▼
10
Publish Data Share anonymized data and documentation; pick repository and access level (fully open is best).
Back to top