Summary of the Anonymization Workflow
TipLearning Objective
- Understand the anonymization workflow.
This overview summarizes the full anonymization workflow. Use it as a reference or checklist when you work through your own data. Keep in mind that not every dataset needs every step in full; how much effort is appropriate depends on your data’s risk (see Does It Always Take This Much Work?).
1
Implement Data Protection
Privacy by design before and during collection: minimize data, plan storage/access, write a Data Management Plan.
▼
2
Collect Data & Handle Identifiers
Use informed consent and pseudonymize early. Remove direct identifiers from every copy, store the pseudonym key separately, and work from a secure identifier-free copy.
▼
3
Run Analysis
Conduct your analysis, writing, and collaborator sharing on the identifier-free working copy.
▼
4
Analyze Attack Scenarios
Consider adversaries and disclosure risks (identity, attribute, inference, membership); set a k-anonymity goal.
▼
5
Calculate Disclosure Risk & Utility
Measure k-anonymity and risk metrics; establish a utility baseline (e.g., using sdcMicro).
▼
6
Choose Anonymization Measures
Select a technique:
non-perturbative,
perturbative,
de-associative, or
synthetic. When in doubt, aim for k-anonymity.
▼
7
Apply Anonymization Techniques
Apply chosen techniques step by step; track changes with sdcMicro.
▼
8
Recalculate Risk, Utility & Statistics
Re-measure after each step. If the balance is not satisfactory,
return to Step 4 and iterate.
▼
9
Document the Process
Internal (auditing) and external (data dictionary) documentation; review scripts; state in README that results may not be fully reproducible.
▼
10
Publish Data
Share anonymized data and documentation; pick repository and access level (fully open is best).