• Home
  • About
  • People
    • Management
    • Staff
    • Scientific Board
    • Advisors
    • Members
    • Fellows
  • Partners
    • Institutional Members
    • Local Open Science Initiatives
    • Other LMU Support Services
    • External Partners
    • Funders
  • Training
  • Events
  1. Research Cycle Handbook

Welcome to the new LMU Open Science Center website! Notice something out of date or incorrect? Let us know!

  • Training Tracks
    • Self-Learning Catalog
      • Principles
        • Assessing Research Replicability
        • Credible Science
        • Replicability Crisis
      • Study Planning
      • Data Management
        • Maintaining Privacy with Open Data
        • Introduction to Open Data
      • Reproducible Processes
        • Advanced Git
        • Readable Code
        • Reproducible Protocols
      • Publishing Outputs
        • Open Access, Preprints, Postprints
    • Research Cycle Handbook
      • Plan & Design
      • Collect & Manage
      • Analyze & Collaborate
      • Preserve & Share
    • Educator Toolkit

On this page

  • 1 Plan & Design
    • Plan & Design Checklist
  • 2 Collect & Manage
    • Collect & Manage Checklist
  • 3 Analyze & Collaborate
    • Analyze & Collaborate Checklist
  • 4 Preserve & Share
    • Preserve & Share Checklist

Other Formats

  • MS Word

Research Cycle Handbook

What is this handbook?

This open research cycle handbook provides detailed practical guidance for making research open, responsible, and reproducible by design. It covers the full research life cycle, from formulating questions and preregistering studies to publishing articles, data, and code. It links to our self-paced tutorials, LMU support services, and essential tools to support open research in everyday practice.

Who is it for?

This manual can be used by individual researchers in any scientific field but is primarily designed to be adapted and adopted by research groups as a whole. Each section outlines concrete actions for ongoing projects and concludes with a team checkpoint.

Navigate our open research cycle handbook
4. Preserve & Share3. Analyze & Collaborate2. Collect & Manage1. Plan & Design
How to build your own lab handbook?
Based on our material, it is easy to build a tailored research practice handbook for your own research group (see example):
  • review the core discipline-agnostic sections by clicking on one of the quadrants in the figure above
  • consult discipline-specific guidance (in development - ETA end 2026)
  • tailor the overview checklist figure
  • use our lab-handbook template as a starting place (in development - ETA end 2026)
You can request a consultation with the LMU Open Science Center: ranging from a 1h one-on-one consultation up to a 6-month training and consultation program for your entire research group (see About this Project below).

This work is part of the “Switch-to-Open Program” (SwOP) funded by the Volkswagen Foundation.

This program consists of a 6-month structured development program for research groups, centered on the co-creation of a tailored research practices lab handbook and culminating in a departmental seminar to facilitate adoption by other groups.

This program is designed to:

  • improve workflow efficiency and reproducibility by establishing clear, standardized procedures and documentation
  • facilitate onboarding and collaboration through a shared, practice-oriented lab handbook
  • support compliance with FAIR and Open Science requirements (see e.g. LMU good research practice guidelines ) and enable researchers to meet evolving recruitment and funding expectations
Interested in entering the Switch-to-Open Program (SwOP) with your research group?
Contact us
How to cite our handbook

Our handbook and other material are licensed CC-BY-SA 4.0. Please feel free to reuse, adapt, and share openly by citing
Ihle Malika, Gupta Reema, Schönbrodt Felix, April 2026, Open Research Cycle Handbook https://lmu-osc.github.io/training/research-cycle-handbook.html CC-BY-SA 4.0 LMU Open Science Center

Acknowledgements

The following people contributed to the writing and revisions of this handbook: Alberto Villagran Asiares, Anja Betz, Pat Callahan, Sara Lil Middleton, Sarah Von Grebmer Zu Wolfsthurn.

1. Plan & Design
  • Explore / reuse research outputs
  • Check legal frameworks
  • Write Research Data Management (RDM) plan
  • Design study and plan statistical analyses

Team checkpoints: study plan presentation + preregistration submission

Go to Handbook - Plan & Design
2. Collect & Manage
  • Collect data following standard protocols
  • Manage data efficiently, and FAIR-ly by design
  • Anonymise data
  • Control data quality

Team checkpoint: data quality validation meeting

Go to Handbook - Collect & Manage
3. Analyze & Collaborate
  • Process and analyze data reproducibly
  • Write reproducible reports and follow reporting guidelines

Team checkpoints: repository & code peer-review + result presentation

Go to Handbook - Analyze & Collaborate
4. Preserve & Share
  • Publish data and/or metadata and data use agreement
  • Publish code (with real, simulated or synthetic data)
  • Deposit a preprint

each with contributorships, persistent identifiers, license.


Team checkpoint: manuscript submission with DOIs for all materials

Go to Handbook - Preserve & Share

1 Plan & Design

4. Preserve & Share3. Analyze & Collaborate2. Collect & Manage1. Plan & Design

Set the foundation for open & reliable research

Explore & Reuse Check Legal Frameworks Write Data Management Plan Design Study
Checkpoints: Study Plan Presentation & Preregistration Submission

1.1 Explore & Reuse

Any resource that inspires you or that you want to reuse and/or adapt must minimally be cited using their persistent identifiers e.g. DOIs (Digital Object Identifiers) - and otherwise URL with author, date, and time of access - and you must follow the license and/or usage agreement provided by the authors. A research output (e.g. data, code) without a license or statement granting your permission for reuse cannot legally be redistributed, and reusing it in work you publish yourself requires the authors’ agreement, even if it appears publicly online.

1.1.1 Articles

  • Review existing literature to come up with a well-founded research question. We recommend to use open source discipline-agnostic registries like OpenAlex which contains published articles, thesis, and preprints (scholarly work that are not (yet) peer-reviewed) of all disciplines, or discipline-specific open source registries such as Europe PubMed Central for life sciences preprints and published articles.

  • Use a reference manager to keep track of your bibliography. Zotero is an open source software formatting your bibliography in any desired format and that can be integrated within e.g. Microsoft Word, Google Doc, or RStudio for writing reproducible manuscripts see (3. Analyze & Collaborate).

LEARN MORE

LMU OSC logo
Zotero logo
OSC Tutorial

Introduction to Zotero

Use an open source reference manager. (1h)

TOOLS & RESOURCES

OpenAlex logo

OpenAlex

All the world's research, connected and open.

Europe PMC logo

Europe PubMed Central

Comprehensive access to life sciences literature.

1.1.2 Preregistrations

A preregistration typically consists of a hypothesis and predictions, a plan for data collection (when relevant), and a plan for data analysis, that researchers upload onto a registry before starting their projects, often in order to increase the rigor of confirmatory research (see 1.4.1. Pre-analysis planning).

  • Get insight into projects that are not (yet) published, either currently ongoing or abandoned, by looking for projects that were preregistered. Projects that are left unpublished typically have a note attached to their preregistration. Some registries are discipline-specific while others are discipline-agnostic (see below).

TOOLS & RESOURCES

OSF icon

Open Science Framework

Registry of preregistrations. Widely used across fields.

AsPredicted icon

AsPredicted

Registry of simple preregistrations

PreclinicalTrials.eu icon

PreclinicalTrials.eu

Registry of preclinical animal study protocols.

animalstudyregistry.org icon

AnimalStudyRegistry.org

Registry of animal studies.

ClinicalTrials.gov icon

ClinicalTrials.gov

Registry of clinical trial protocols.

1.1.3 Data

How to find existing datasets?
  • Search for discipline-specific repositories on re3data which is a central registry of many repositories
  • Explore subject agnostic repositories such as DataCite, FigShare, Open Science Framework (OSF), or Zenodo.

These platforms either give you access to existing data or provide metadata and explanations on how to request access to the data.

NoteDefinition

Metadata are data about your data, such as author, date, measurement device, unit of measurement, context of data collection, etc.

How to reuse a dataset?
  • Review the license and data use agreement. Make sure you understand what you are allowed to do with the data and under what conditions. Even if the license does not request attribution of the authors, scholarly norms require you to cite the source of the data for any of your work based on it.
  • Review metadata and documentation. Make sure you know where your data comes from, how the data was collected and processed, and reflect on whether any of it poses problems for your research question.
  • Check what additional requirements the data sources have. Sometimes, data providers request prospective data users to submit a preregistration prior to giving access to the data (see 1.4. Study Design & Analysis Plan).
  • Use the metadata to plan your analysis. Review existing data dictionaries (or “codebooks”) and other documentation describing the variables, range of values, etc. If you plan to do a confirmatory analysis, do not look at the data to minimize confirmation or hindsight bias; instead, prepare a pre-analysis plan (see 1.4. Study Design & Analysis Plan).
NoteDefinition

Confirmation bias: The tendency to seek, interpret, and remember information that confirms one’s existing beliefs or expectations.

Hindsight bias Seeing past events as predictable after the outcome is known (“I knew it all along”).

TOOLS & RESOURCES

re3data icon

re3data

Registry of research data repositories.

DataCite icon

DataCite Commons

Discovery tool connecting works, people, and organizations.

FigShare icon

Fig Share

General-purpose repository for data, software, reports.

OSF icon

Open Science Framework

General-purpose repository for data, materials, reports.

Zenodo icon

Zenodo

General-purpose repository for data, software, reports.

1.1.4 Code

Find code available for reuse archived on Zenodo, Software Heritage or actively developed on GitHub and other code repositories. Start learning Git version control now or learn to take advantage of more collaborative features on the GitHub platform in more details in 3. Analyze & Collaborate.

Important

Code publicly visible on GitHub without a license or equivalent text explicitly stating permission for reuse cannot legally be redistributed, and reusing it in work you publish yourself requires the authors’ permission. It is best to ask the authors to add an open license to their repository to explicitly allow reuse (to do this, they can, for instance, add a file called LICENSE.txt with the Apache 2.0 license text - see our code publishing tutorial to learn more about licenses).

LEARN MORE

LMU OSC logo
Git logo
OSC Tutorial

Git Tutorial

Use version control system Git from within RStudio. (2h)

LMU OSC logo
GitHub logo
OSC Tutorial

GitHub Tutorial

Collaborative coding with GitHub and RStudio (1h)

LMU OSC logo
OSC Tutorial

Code Publishing

Add README and license to a reproducible project (2h)

TOOLS & RESOURCES

Zenodo icon

Zenodo

General-purpose repository for data, software, reports.

Software Heritage icon

Software Heritage

Collects, preserves, and shares software in source code form.

GitHub icon

GitHub

Cloud-based platform to collaborate on code.

1.2 Legal Requirements

1.2.1 LMU guidelines

The LMU Guidelines for Safeguarding Good Scientific Practice are legally binding for all academics, researchers, research support staff, teachers, and students at LMU Munich. Only the original text in German prevails, but we provide an English summary of relevant aspects for this guide:

Appropriate level of documentation and standards to allow reproduction:
  • Reproducible methods must be used. (§11)
  • When research software is developed, its source code must be documented. (§12)
Appropriate level of documentation and standards to allow replication:
  • All information relevant to the production of a research result must be documented comprehensively to enable replication. (§7 and §12)
  • If specific professional recommendations exist for review and evaluation, the results must be documented in accordance with these respective specifications. (§12)
  • Individual results that do not support the hypothesis must also be documented; a selection of results is not permitted. (§12)
Public access to research results:
  • Apart from specific exceptions, all findings should be made public. For this, they must be described in a detailed and comprehensible manner which includes making available the research data, materials and information on which the results are based, as well as the methods used and the software employed (including appropriately licensed self-written software) according to the FAIR principles. (§13)
  • Data, material, software made publicly accessible must be appropriately archived, usually for a period of 10 years (§17).

In later sections, you will acquire skills in data management and reproducible workflow that will enable you to comply with these guidelines and the FAIR principles.

NoteDefinitions

The FAIR principles are defined as:

  • Findable: metadata should be deposited in a searchable repository and be assigned a permanent identifier
  • Accessible: the data is either open, or accessible upon some authentication process, or closed, but with open metadata.
  • Interoperable: the data is described with a standard terminology (so the dataset can be merged with other ones) and saved in a stable file format
  • Reusable: the data is richly documented (e.g. with a data dictionary) and is accompanied by a data usage license See https://www.go-fair.org/fair-principles/ for more information.

Metadata are data about your data, such as author, date, measurement device, unit of measurement, context of data collection, etc.

Reproducibility: The ability of a researcher to re-derive the same results using the same data and methods; also known as computational reproducibility.

Replicability: The ability of an independent researcher to achieve results consistent with the original study by following the same experimental or analytical approach but collecting new data.

TOOLS & RESOURCES

LMU logo

LMU Guidelines for Safeguarding Good Scientific Practice

Implementation of the German Research Foundation's (DFG) Code of Conduct

1.2.2 Funders

  • Check all funders’ open science requirements. Funders may have additional requirements on top of those indicated in the LMU guidelines. For instance, some funding lines request a Research Data Management plan before making their second payment, some specify the extent and timing of data sharing and provide funds for such activity.

  • Contact the LMU Research Funding Unit to review your grant proposal and assess if your proposal is meeting your funders’ open science requirements.

1.2.3 Ethics

Data collection and analyses involving human participants or animal subjects typically require approval from ethics committees to ensure responsible conduct and the protection of data.

Your ethics proposal will typically include information on:
  • Data storage and retention – outlining how data will be securely stored, backed up, and retained over time. This information can be extracted from a more detailed Research Data Management plan (see 1.3. Research Data Management Plans).
  • Risks if the data were leaked – identifying potential consequences for participants or the research project if confidentiality is breached.
  • Data anonymization – describing procedures to remove or obscure personally identifiable information to protect participant privacy (see 2.3.2. Anonymization for options, from simple techniques of anonymization to the creation of synthetic data).
  • Informed consent forms language – ensuring that participants clearly understand the purpose, procedures, and any potential risks of the study. Conditions for sharing their data should be clearly explained here (see 2.3.1. Informed Consent).
  • Power analysis to justify sample size – providing a statistical rationale for the number of participants, which supports the validity and ethical justification of the study. This, and more detailed information on the statistical plan, can be extracted from your pre-analysis plan (see 1.4.1. Pre-analysis planning and 1.4.3. Power analyses).

For data protection guidance, contact the LMU Data Protection Officer or the Research Data Management team of the University Library.

TipTips for research groups to streamline this process
  • Share templates and example resources amongst team members. For example, include previously approved ethics proposals, approved Data Protection Impact Assessment forms, or Data Management Plans on a common server space.
  • Create Standard Operating Procedures (SOPs) for the team for processes such as appropriate anonymization technique for a specific data type, power analyses script for common analyses, define when a data management plan must be updated, who is responsible, and how updates are reviewed/approved and communicated.

LEARN MORE

LMU OSC logo
OSC Tutorial

Data Management Plans

Overview of components, tips, and tools. (30 min)

LMU OSC logo
OSC Tutorial

Data Anonymization

Implement data anonymization techniques in R. (3h)

LMU OSC logo
OSC Tutorial

Power Analysis

Data simulations for GLMs, LMEs, and SEMs in R. (6h)

1.3 Research Data Management Plans

A Data Management Plan (DMP) documents how you will handle research data throughout your project. Writing a DMP prompts you to think and document decisions you might otherwise leave implicit.

  • Decide before data collection whether you will eventually share your data publicly (and where), in order to (i) get ethics approval on the right plan, (ii) design consent forms for participants, (iii) collect appropriate metadata for the target repository, etc.
  • Start with what you know, and refine the details as your project develops. Your DMP is a living document that you will refine to match the reality of your project while ensuring data protection and streamline collaborations (see 2.2. Data Management, 3.1. Data Processing & Analysis, and 4.1. FAIR Data Sharing).

Your DMP will ask:

  • What data will you collect or generate (types, formats, volume, sources)? See 2.1. Data Collection.
  • How will you describe it (metadata standards, documentation practices)? See 2.2. Data Management for these and the next questions.
  • How will you organize files (naming conventions, folder structure, versioning)?
  • Where will you store it (locations, backups, access controls)?
  • How will you ensure quality (validation checks, error-handling)?
  • How will you share outputs (repositories, licenses, embargo periods)? See our lecture “Why share data openly?” and 4.1. FAIR Data Sharing
  • What constraints apply (consent, anonymization, GDPR, data use agreements)? See our lecture “Maintaining privacy with open data”, 1.2.3. Ethics and 2.3. Ethics & Privacy.

The specific questions vary by discipline, data type, and funder requirements. DMP tools like RDMO guide you through the relevant questions with funder-specific templates.

TipTips for research groups to streamline this process
  • Share templates and example DMP amongst team members on a common server space.
  • Create Standard Operating Procedures for the team. Define when a data management plan must be updated, who is responsible, and how updates are reviewed/approved and communicated.

LEARN MORE

LMU OSC logo
OSC Lecture

Why share data openly?

An introduction to the what, why, and how to make data open (30 min)

LMU OSC logo
OSC Lecture

Maintaining Privacy with Open Data

How to make data open without revealing sensitive information (1h)

LMU OSC logo
OSC Tutorial

Data Management Plans

Overview of components, tips, and tools. (30 min)

TOOLS & RESOURCES

LMU OSC logo
RDMO logo
Supported at LMU

RDMO

Funder-compliant DMP templates (e.g. DFG, ERC).

RIOjournal icon

RIOjournal

Examples of DMPs by discipline.

1.4 Study Design & Analysis Plan

1.4.1 Pre-analysis planning

Why should you plan your statistical plan prior to collecting data?

Humans are prone to cognitive biases such as confirmation bias (seeking information that supports existing beliefs) and hindsight bias (believing outcomes were predictable after the fact). In research, these biases can distort findings, especially when researchers make analytic decisions after seeing results. Although statistical testing typically accepts a 5% false positive rate, “researcher degrees of freedom” — choices about data collection, exclusions, transformations, sample size, covariates, etc. — can dramatically inflate false positives when decisions are made post hoc. Practices like increasing sample size until reaching statistical significance, selectively removing outliers, or trying multiple analytic strategies increase the likelihood of false-positive results. See how easy it is to find false “significant” results by using our p-hacking tool.

The core problem is that analyses guided by observed outcomes allow biases to influence decisions, making many reported effects unreliable. A key remedy is transparency and preregistration.

Benefits of preregistration

Preregistration, that is, specifying hypotheses, methods, and analysis plans before data collection or analysis, limits bias in confirmatory testing while still allowing exploratory analyses, clearly distinguishing robust hypothesis tests from hypothesis-generating work. This improves credibility, limits false positives, and often leads to better study design through early methodological feedback.

Preregistration can be beneficial for various type of studies, including:

  • experimental studies (i.e. studies with a manipulated variable): it will define what will be your confirmatory analysis and strengthens your claim
  • observational or exploratory studies: it will help you move along the exploratory-confirmatory continuum
  • qualitative studies: it will provide a way to document e.g. your positionality towards a subject in the course of a project.
What is included in a preregistration?

Several preregistration templates exist. While the standard Open Science Framework (OSF) preregistration template is most commonly used, some are tailored for specific field or specific methods (e.g. systematic review, qualitative work, secondary data analysis).

Your preregistration will define your study’s:

  • Hypothesis and predictions
  • Data collection procedures
  • Sample size and stopping rule
  • Variables (manipulated, measured, indices)
  • Statistical method (model, dependent and independent variables, covariables, transformations)
  • Data exclusion criteria
  • How to deal with missing data

A great tool to create your statistical plan, especially for early career researchers still learning statistics and needing feedback from supervisors, collaborators, or statisticians on their design, is to simulate data, and write the possible statistical tests to analyze that data (see 1.4.2. Simulation of data and 1.4.3. Power analyses). Including an analyses script (developed on simulated data) with your preregistration is optional but recommended.

To get support with planning your analytical approach, you can book a consultation with the LMU statistical consulting unit (StaBLab).

Publishing process

Once your study plan is finalized:

  • Submit your preregistration before collecting new or analyzing existing data. You can do so on discipline specific registries (see 1.1.2. Preregistrations) or discipline agnostic repositories such as the OSF.
  • Embargo your plan if you are concerned about scooping. On the OSF, your preregistration can be kept private for a predetermined amount of time, and for a maximum of 4 years.
  • Include your preregistration’s DOI in your manuscript. Make your registration public upon the publication of your manuscript.

Creating a preregistration improves transparency and allows for valuable early feedback from collaborators. An even stronger approach is submitting preregistrations directly to journals (then called “Registered Reports”), enabling peer review at a stage where methodological adjustments are still possible.

Registered Reports

Registered Reports are a publication format, now adopted by over 300 journals (see participating journals), where preregistrations are peer-reviewed before data collection. Reviewers evaluate the hypotheses, methods, and planned analyses, allowing methodological improvements. If the plan is approved, the journal grants in-principle acceptance, meaning publication is guaranteed provided researchers follow the protocol.

After completing the study, authors add results and discussion sections, clearly separating preregistered confirmatory analyses from exploratory ones. Final review focuses on adherence to the approved plan and the validity of conclusions, not on whether results are significant. This model shifts incentives toward asking important questions and using rigorous methods rather than chasing striking or ‘novel’ outcomes.

LEARN MORE

LMU OSC logo
OSC Tutorial

TBA: Preregistration tutorial

Step-by-step guide to creating preregistration. (Xh)

TOOLS & RESOURCES

LMU OSC logo
OSC Tool

P-hacking tool

Interactive app to realize how easy it is to find false "significant" results.

COS icon

Center for Open Science

List of journals offering Registered Reports.

OSF icon

Open Science Framework

Preregistration templates, embargoes, file storage.

1.4.2 Simulation of Data

In our context, a computer simulation is the generation of artificial data to build up an understanding of real data and the statistical models we use to analyze them. You can simulate data to:

  • Test your statistical intuition or demonstrate mathematical properties you cannot easily anticipate.
    Example: Check whether there are more than 5% significant effects (assuming \(\alpha = .05\)) when random data from \(H_0\) are generated.

  • Understand sampling theory and probability distributions or test whether you understand the underlying processes of your system.
    Example: See whether simulated data drawn from specific distributions is comparable to real data.

  • Perform power analyses.
    Example: Assess whether the sample size (within a simulation repetition) is high enough to detect a simulated effect in more than 80% of the cases. (see 1.4.3. Power analyses)

  • Prepare a pre-analysis plan.
    Example: To strengthen your planned confirmatory analyses before collecting data, consider sharing a simulated dataset with a statistician or mentor. This allows for specific feedback on suitable statistical tests. The resulting analysis code can accompany your preregistration or registered report (see 1.4.1. Pre-analysis planning) so reviewers can clearly see your intended approach. When real data are collected, they can be directly substituted into the code to generate results.

Generating an artificial dataset in R (see our simulation tutorial) is much easier than you might think and is often very helpful, even when you need to make assumptions about variable distribution or when the parameter space is not well known.

LEARN MORE

LMU OSC logo
R logo
OSC Tutorial

R Tutorial

Learn R programming. (3h)

LMU OSC logo
OSC Tutorial

Data simulation in R

Easy data simulations in R. (2h)

1.4.3 Power analyses

Power analysis is relevant whether you are designing a project from scratch or running an analysis on already existing data. There are two main types of power analyses:

A priori power analysis

Simulate data to calculate the smallest sample size required to detect the smallest effect of interest. See our advanced power analyses tutorial using R.

For a very basic power calculation, you can use simple R functions if you know 3 out of 4 of these parameters:

  • required sample size n (usually the one missing)
  • desired power 1 - β (default 0.80)
  • the alpha level α (default 0.05)
  • the expected effect size (has to be estimated or extracted from the literature on the form of d, f, etc.)

To get support with pre-analysis planning, you can book a consultation with the LMU statistical consulting unit (StaBLab).

Post-hoc power analysis

Compute a post-hoc power when you are not be able to control the sample size for your project. Beware: This power computation comes in two flavors - one is legitimate, and one is flawed and not defensible.

The legitimate post-hoc power is computed with your actual n, and the same effect size that you plugged into your a-priori power analysis. This analysis gives you the achieved power to detect your assumed effect.

The flawed version of post-hoc power is called “observed power”: If an analysis yields a non-significant result, some researchers calculate the post-hoc power, but plug in the observed effect size. “Observed power”, however, is just a one‑to‑one function of the p‑value (a non-significant p-value returns a low power < 50 %, a just significant p‑value of .05 always yields a power of exactly 50%). Observed power adds no new information to the p‑value and is essentially meaningless. Do not compute this type of post-hoc power!

LEARN MORE

LMU OSC logo
R logo
OSC Tutorial

R Tutorial

Learn R programming. (3h)

LMU OSC logo
OSC Tutorial

Power Analyses

Data simulations for GLMs, LMEs, and SEMs in R. (6h)

Plan & Design Checklist

To complete before presenting your final study plan to your research group and, if applicable, submitting your ethics proposal and/or preregistration. Not all items are relevant for all fields of research or study types.

Background Information

Study Design

Data Management Planning

Project Management

Before Data Collection

Download checklist

2 Collect & Manage

4. Preserve & Share3. Analyze & Collaborate2. Collect & Manage1. Plan & Design

Build quality into your data from the start

Data Collection Data Management Ethics & Privacy Quality Control
Checkpoint: Data quality validation meeting

2.1 Data Collection

Your data acquisition procedures must be documented in sufficient detail to allow replication by another researcher (see LMU Guidelines for Safeguarding Good Scientific Practice). Reproducible data collection processes build team expertise, reduce errors, and improve data quality and consistency.

State-of-the-art practices for reproducible data acquisition include:

  • creating standard operating procedures
  • recording metadata as data collection is taking place
  • build in automation through programming
NoteDefinition

Metadata are data about data; they provide context to your data. Metadata such as equipment settings, environmental conditions, software versions, and calibration records should be recorded contemporaneously, not reconstructed afterward. Electronic lab notebooks, instrument logs, and automated logging all help to document your metadata.

TipTips for research groups to streamline this process
  • Create standard data acquisition procedures within the team. From step-by-step wet-lab procedure, to the settings of measuring devices and the reproducible data pre-processing using script, all regularly repeated steps should be documented and standardized to be replicated precisely by all team members.
  • Share standard operating procedures through common server space such as LRZ Sync & Share, specialized online tools like protocols.io or electronic lab notebooks, or LRZ GitLab for scripts.

2.1.1 Lab Protocols

Your protocol should specify materials with identifying details (e.g. lot numbers, versions, sources), equipment settings, step-by-step instructions with timing, and expected outcomes at each stage. What counts as “materials” varies by field: reagent concentrations in wet lab work, scanner parameters in neuroimaging, sampling coordinates in field ecology. But the principle is the same: enough detail that someone else could replicate your procedure exactly.

  • Write detailed methods and reusable protocols. Write your protocol before you start, with all the details that would be needed for an exact replication (see the ReproducibiliTeach lecture on reusable protocols)

  • Track deviations in real time. Follow your protocol precisely, and record any deviations as they happen. When you need to adapt, note it immediately. These deviations often explain unexpected results and guide protocol improvements. Electronic lab notebooks (ELNs) make this easier by creating version-controlled, timestamped records automatically, providing an audit trail that paper cannot match.

  • Publish your protocols. A detailed, tested protocol is a contribution to your field. Publishing establishes priority, enables citation, and makes your methods reusable. Platforms like protocols.io provide version control and DOI assignment.

LEARN MORE

reproducibiliteach icon

Write reusable protocols

Understand the level of details needed for your protocols.

TOOLS & RESOURCES

eLabFTW icon
Supported at LMU

eLabFTW

Electronic Lab Notebook hosted by LMU Munich.

Chemotion icon
Supported at LMU

Chemotion

ELN hosted by the Faculty for Chemistry and Pharmacy.

Protocols.io icon

Protocols.io

Share, discover, cite, and improve research protocols.

2.1.2 Questionnaires

Document both the instrument and the administration procedure completely. The data dictionary (or ‘codebook’) should record the exact version of the questionnaire used; if copyright allows it, document the item wording itself.

Questionnaires should be objective, reliable, and valid:

  • Objective: Results should not depend on who administers or scores the questionnaire.
  • Reliable: Responses should be consistent across repeated measurements when the underlying construct is unchanged (i.e., they should have low measurement error)
  • Valid: Items should measure the intended construct rather than something else. Typical ways to assess validity include content validity, convergent and divergent construct validity, and criterion-related validity.

For well-validated instruments, published studies extensively assess reliability and validity across several populations and contexts. If these quality criteria of measurement instruments are met, your effect size and statistical power will be increased.

2.1.3 Software-based data acquisition

When data comes from instruments, sensors, or APIs (Application Programming Interface), scripting the acquisition creates a reproducible record of exactly what was collected and how. Programming languages like R or Python work well for straightforward pipelines. For more complex multi-step workflows which are common in e.g. in bioinformatics and neuroimaging, workflow managers like Snakemake ensure steps run in the correct order and can resume after failures.

  • Structure data correctly from the start. Variables in columns, observations in rows. This makes your data immediately interoperable with analysis tools rather than requiring cleanup later. Scripts can also automate organization, file renaming, and conversion to open formats. See 2.2 Data Management for guidelines.

  • Keep records of what ran and when. Include error handling so failures are recorded rather than silently corrupting data. When something fails months later, you need to know what happened. Always test acquisition scripts on sample data before production runs. A bug in your collection pipeline can invalidate an entire dataset.

  • Version control your code and data. This makes your methods reproducible and shareable. See 2.2.6. Version Control for details.

LEARN MORE

LMU OSC logo
R logo
OSC Tutorial

Introduction to R

Programming fundamentals for research data processing (3h)

LMU OSC logo
Git logo
OSC Tutorial

Introduction to Git

Learn Git basics integrated with RStudio (2h).

2.1.4 Field and lab work

Field and sometimes lab researchers have to work with analogue notebooks first, and use temporary storage solution for images and video-recordings.

  • Digitalize your data soon after data collection, e.g. using a data entry form from a relational database such as PostgreSQL or SQLite. Any relational database would work, and Excel/CSV files are also common options
  • Match the tool to the size of your project. A relational database pays off for ongoing, multi-site, or long-term data entry, but for small or one-off projects a structured spreadsheet template is enough (see 2.2.3. File Formats for the pitfalls to avoid)
  • Create automatic validation checks for data entry so e.g. values outside of expected range get immediately identified (see 2.4. Quality control)
  • Transfer digital data (e.g. pictures, video recordings) from temporary storage solutions (e.g. camera storage) to permanent storage solutions (see 2.1.1. Storage)
  • Check your data entries, e.g. after a day, to correct possible mistakes in transcription
  • Keep analogue records to correct data entries errors at the end of the season or experiment

TOOLS & RESOURCES

PostgreSQL icon

PostgreSQL

Open source relational database software

2.1.5 Systematic reviews

A high-quality systematic review, whether it includes a meta-analysis or not, uses a rigorous, transparent, and reproducible methodology to select the relevant literature and later provide a thorough summary and critical evaluation of research within a given field. Its defining features include:

  • clearly defined objectives supported by an explicit and reproducible methodology
  • a comprehensive, systematic search designed to identify all studies meeting the eligibility criteria
  • a critical appraisal of the validity of included studies, such as through risk-of-bias assessment
  • a structured presentation and synthesis of the characteristics and findings of the included studies

Data collection, in this context, consists in developing a judicious reproducible methods to selecting relevant literature.

  • Use tools such as

    • CAMARADES which provide a supporting framework for the conduct of systematic reviews of animal studies, and
    • PRISMA (Preferred Reporting Items for Systematic review and Meta-Analysis Protocols) and its extensions such as PRISMA-P which offers checklist for systematic review protocol reporting.
  • Learn more from our colleagues at the Berlin Institute of Health QUEST Center for Responsible Research using their resources for systematic reviews in biomedical research and the material of their systematic review workshop at our previous OSC summer school.

LEARN MORE

LMU OSC logo
OSC Workshop

Systematic review

Slides and additional resources of workshop delivered at OSSS23.

TOOLS & RESOURCES

PRISMA icon

PRISMA

Preferred Reporting Items for Systematic review and Meta-Analysis Protocols

CAMARADES icon

CAMARADES

Step-by-step guide to Preclinical Systematic Review.

2.2 Data Management

In 1. Plan & Design you created a Research Data Management Plan. It’s now time to put this plan into practice, refining it as you learn what actually works for your project.

  • Keep your raw data as read-only files. The unmodified output of your instruments, surveys, or observations should never be modified directly. You should provide enough documentation to recreate your processed datasets and results from your raw data (see 3.1. Data Processing & Analyses).

  • Organize, document, and store your research files so they remain usable, and become FAIR upon sharing. Beyond raw data, you will generate processed data, code, documentation, and metadata. How you organize, describe, and store these determines whether your work remains usable and reproducible. The FAIR principles guide these decisions: making outputs Findable, Accessible, Interoperable, and Reusable.

What follows are general practices. Your domain has specific conventions for file formats, folder structure, and metadata. RDMkit provides detailed guidance organized by research area.

TipTips for research groups to streamline this process
  • Create standard operating procedures (SOPs) within the team. For instance: where and when data back up should be made and in which file format, what project folder structure and conventions for naming files should be adopted, which metadata should routinely be acquired, what documentation should be created and when, and how and where a history of versions should be preserved.
  • Assign responsibilities. For developing such SOPs for a specific kind of project and for checking all members have implemented those SOPs at a specified time point in their project (e.g. hold a data quality validation meeting prior to starting data analyses, see Collect & Manage Checklist).

2.2.1 Storage

This section covers storage for data you are actively collecting. Long-term archiving for sharing is covered in 4. Preserve & Share.

  • Use institutional storage. LMU Munich provides storage such as LRZ Sync and Share or LRZ DSS with automated backups, access controls, and GDPR compliance. Additional options vary by department. Contact the Research Data Management team of the University Library to find what is available to you. When choosing, consider how much data you will generate, who needs access, and whether your data includes personal information requiring stricter controls.
  • Follow the 3-2-1 backup rule. Keep three copies on two media types with one off-site. Designate one location as the master copy, the authoritative version everything else syncs from. Working with multiple “equal” copies creates version conflicts. Remember that syncing is not backup: if you delete a file from a synced folder, the deletion propagates everywhere. True backups preserve previous versions independently.
  • Control access from the start. Grant access only to those who need it. Use institutional sharing tools, not email attachments or personal cloud links. For collaborations, agree at the start who can read, who can edit, and who manages permissions. When team members leave, remove their access promptly.
  • Test your backups. A backup you cannot restore is not a backup. Test restoration at least once. Archive inactive data periodically and review access lists when team composition changes.
ImportantAvoid for Research Data

Personal laptops as primary storage, external drives as only copy, consumer cloud services (Dropbox, Google Drive) for sensitive data, and USB drives except for temporary transport.

LEARN MORE

LRZ logo
Supported at LMU

LRZ Sync & Share

Cloud storage service for LMU Munich

LRZ logo
Supported at LMU

LRZ DSS

Long-term archival storage for LMU Munich

OSF logo

OSF

Research project management platform including storage.

2.2.2 Organization

Your folder structure and file naming conventions determine whether you and others can navigate your project months or years later. Establish these conventions at the start of your project and document them. When collaborating, ensure everyone follows the same system.

  • Separate raw from processed data. Raw data should not be modified: once collected, these files should never be touched. All cleaning, data processing, transformations, and analyses happen on copies or through scripts in a separate folder. This preserves your ability to verify results or reprocess from the original source.

  • Develop a file naming convention. Good file names identify contents at a glance and sort correctly. Balance specificity with readability: too many elements make names unwieldy, too few make them ambiguous. Order elements from general to specific.

    • Use underscores or hyphens to separate elements, never spaces or special characters (? ! & * % # @)
    • Use ISO 8601 dates (YYYY-MM-DD) so files sort chronologically
    • Include version numbers with leading zeros (v01, v02) so v10 sorts after v09
    • Use meaningful abbreviations and document what they mean

A pattern like YYYY-MM-DD_project_condition_type_v01.ext places files in chronological order while preserving context. For example, 2024-03-15_sleep-study_control_survey_v02.csv immediately tells you when it was created, which project it belongs to, the experimental condition, data type, and revision. Document your convention in a README file stored next to your data files so collaborators can parse filenames without asking.

  • Follow domain standards where they exist. Many fields have established organizational conventions that tools and collaborators expect. Using these means your data can immediately be integrated in existing analysis pipelines and reviewers recognize the structure. Search RDMkit for standards in your domain.

LEARN MORE

LMU OSC logo
OSC Tutorial

Data Organization

Folder structure and naming conventions. (30 min)

TOOLS & RESOURCES

LMU OSC logo
GitHub logo
OSC Tool

Research Project Template

Project folder separating raw & processed data, code, and outputs.

RDMkit logo

RDMkit

Domain-specific data management standards

2.2.3 File Formats

File format choices affect who can work with your data now and whether it remains readable in the future. Open formats have publicly documented specifications that anyone can implement, so many programs can read them and they remain accessible even if the original software disappears. Proprietary formats lock you into specific tools, complicate collaboration, and risk becoming unreadable if the company stops supporting them.

  • Keep raw data in its original format. Whatever your instrument or source produces, preserve that original as your ground truth. Even if it is proprietary, you need it for verification and potential reprocessing.

  • Work in open formats. For analysis, convert to open formats like CSV, JSON, or plain text. This makes your workflow reproducible, enables collaboration across different tools, and ensures your data can be shared. If conversion loses important information (metadata, precision, structure), document what is lost and keep both versions.

  • Be careful with spreadsheets. Excel is convenient for data entry but causes real problems. It silently converts data: gene names like MARCH1 become dates, leading zeros in IDs disappear, and long numbers lose precision. Formatting (colors, merged cells) breaks machine-readability since scripts cannot see it. If you use spreadsheets for entry, keep them simple (one header row, one observation per row, no merged cells) and export to CSV immediately. Save CSVs with UTF-8 encoding to avoid character corruption when sharing across systems. For more guidance on spreadsheet best practices, see The Turing Way and UC Davis DataLab.

  • Check domain recommendations. Your field likely has established conventions balancing openness with practical needs like performance or metadata preservation. Consult the RDMKit to find conventions for your field.

Format issues often surface during quality control. The 2.4. Quality Control panel below covers validation checks that can catch encoding problems, unexpected conversions, and structural inconsistencies early.

TOOLS & RESOURCES

RDMkit logo

RDMkit

Domain-specific file formats and conventions

2.2.4 Documentation

Without documentation, a dataset is just a collection of files. Six months from now, you will not remember what each column means, why certain values are missing, or how files relate to each other. Documentation makes your data usable by your future self, your collaborators, and any other researchers.

  • Create a README file (as .md or .txt) early and update it as you go. Your README is the entry point to your project. Start it when you begin, not when preparing to publish. A good README answers the essential questions: who created the data, what it contains, when and where it was collected, why it was generated, how it was produced, and whether it can be reused. These answers let someone unfamiliar with your project understand and work with your data.
  • Create a data dictionary defining every variable. A data dictionary (or “codebook”) makes your dataset self-explanatory. For each variable, document what it measures, its data type, valid values, units of measurement, and how missing data is coded. Use appropriate missing codes to distinguish why data is absent (declined to answer, not applicable, technical failure) since this distinction matters for analysis.

LEARN MORE

LMU OSC logo
OSC Tutorial

Principles of Data Documentation

Principles of README files and data dictionaries.

LMU OSC logo
OSC Tutorial

Data Documentation & Validation

Create READMEs, data dictionaries, and validation checks for your data. (1h)

2.2.5 Standards

Standards are community agreements on how to organize and describe research data. Using them means others in your field immediately understand your data. Three types of standards matter here:

  • Organizational standards specify how to structure files and folders. Some fields have well-established conventions, like BIDS for neuroimaging data. When such standards exist, use them. Your data will work immediately with existing tools, and collaborators will recognize the structure without explanation. If no standard exists for your domain, create a consistent structure and document it in your README. See for instance our research project template.

  • Reporting guidelines specify what methodological details to document for different study types. The EQUATOR Network maintains a searchable database of guidelines for clinical trials, observational studies, animal research, and many other study types. Following these ensures you capture everything others need to understand or replicate your work.

  • Metadata standards define what descriptive information to record and how to structure it. Scientific metadata describes how your data was produced: equipment specifications, acquisition parameters, protocols followed. This is distinct from discovery metadata (titles, keywords, descriptions) which you will prepare when sharing in 4. Preserve & Share. Your field has conventions for which parameters matter. FAIRsharing catalogs metadata standards by discipline.

Think of your data as a first-class research output. Comprehensive metadata transforms a project artifact into a reusable resource. Someone reanalyzing your data years later needs to understand exactly how it was produced.

TOOLS & RESOURCES

EQUATOR logo

EQUATOR Network

Comprehensive database of reporting guidelines.

FAIRsharing logo

FAIRsharing

Search by discipline to find metadata standards, reporting guidelines, and data policies for your field.

RDMkit logo

RDMkit

Domain-specific metadata standards

LMU OSC logo
GitHub logo
OSC Tool

Research Project Template

Project folder separating raw & processed data, code, and outputs.

2.2.6 Version Control

Version control system like Git tracks changes to files over time. You can see what changed, when, and why. You can revert to previous versions. Collaborators can work without overwriting each other.

  • Git usually suffices for data files. Text-based formats (CSV, JSON, plain text) and smaller binary files work well in standard Git repositories. You get a complete history of changes and can share easily via GitHub or GitLab.

  • Use specialized tools for large or frequently changing binary files. Standard Git stores each version in full, so repositories become unwieldy with large datasets. Git LFS (Large File Storage) stores large files separately while keeping them tracked. Git-annex manages files across multiple storage locations. DataLad builds on git-annex and works with standard Git workflows.

NoteDifference between Git, GitHub and GitLab
  • Git is a version control system that tracks changes in text files (e.g. CSV, plain text, R, Python). The Git software and your Git repositories should be, respectively, installed and located in your local environment (i.e. on your computer, not on a drive, see Git tutorial).

  • GitHub is the most popular, free but proprietary and US-based cloud-based platform for software development with Git, providing collaboration features like pull requests and issues (see GitHub tutorial). You should not have any sensitive data on GitHub even in a private repository.

  • LRZ GitLab is a cloud-based hosting platform that provides essentially the same features as GitHub but is open source and is installed on the LRZ servers for LMU Munich and can therefore be considered secure when the repository is private.

While your LRZ GitLab account is associated with your LMU Munich affiliation, your GitHub account can be associated with your private email, be included in your CV, and be used for public sharing of your data and code (see 3. Analyze & Collaborate and 4. Preserve & Share).

In a version controlled workflow, you back up your local Git repositories on either GitHub or LRZ GitLab through a secure SSH connection (see GitHub tutorial) and share access to your repositories with your collaborators through the cloud-based platform GitHub or LRZ GitLab.

LEARN MORE

LMU OSC logo
Git logo
OSC Tutorial

Introduction to Git

Learn Git basics integrated with RStudio (2h).

LMU OSC logo
GitHub logo
OSC Tutorial

Introduction to GitHub

Connect to GitHub from Git within RStudio (1h).

TOOLS & RESOURCES

LMU OSC logo
GitLab logo
Supported at LMU

LRZ Gitlab

Institutional Git hosting for LMU Munich.

DataLad icon

DataLad

Version control for large datasets

2.3 Ethics & Privacy

Research involving human participants requires ethics approval and data protection compliance. In your ethics proposal (see 1.2.3. Ethics) you planned for safeguards for the people contributing to your research and you now need to implement them before or while collecting your data.

2.3.1 Informed Consent

Participants have the right to understand what they are agreeing to. Your consent form should explain the research purpose in plain language, describe what data you will collect and how you will protect it, specify who will have access and for how long, and make clear that participation is voluntary. See “How to write an informed consent form” from the University to Utrecht and an example template from the LMU Psychology Department.

  • Use tiered consent when you plan to share data. Some participants may consent to their data being used for your study but not shared publicly. Others may be comfortable with broader sharing. Giving options respects autonomy while maximizing what you can eventually share.

  • Store consent forms separately from data. The consent form links a name to participation. Keeping it with your data undermines any pseudonymization you apply.

TipTips for research groups:

Maintain a centralized log of all ethics approvals, consent forms, and compliance certifications for easy reference by team members.

TOOLS & RESOURCES

University of Utrecht icon

How to write an informed consent form

University of Utrecht RDM guide.

2.3.2 Anonymization

Anonymization protects privacy and determines what you can share.

  • Remove direct identifiers during collection. Names, addresses, ID numbers, photographs, email addresses. Replace these with codes.

  • Assess indirect identifiers carefully. A combination of age, location, profession, and a rare condition might identify someone even without their name. Timestamps reveal patterns. Free-text responses often contain identifying details participants did not intend to share. Follow our data anonymization tutorial to learn to evaluate and implement anonymization techniques in R.

  • Generate synthetic data when full anonymization or real data sharing is not possible. Synthetic data is artificially generated data that can serve as a privacy-preserving alternative to sensitive datasets, enabling researchers to reproduce analyses, verify findings, and initiate model development when access to real data is restricted. Sharing synthetic data together with code, rather than sharing no data at all, increases the utility of your research (see 4. Preserve & Share).

NoteDistinction between pseudonymization and anonymization
  • Pseudonymization replaces identifiers with codes while retaining a key that links back to individuals. Pseudonymized data is still personal data under GDPR because re-identification is possible, at least by some persons.

  • Anonymization removes all possibility of re-identification. Only truly anonymized data falls outside GDPR scope. Achieving this is harder than it appears, especially with rich datasets.

LEARN MORE

LMU OSC logo
OSC Tutorial

Data Anonymization

Implement data anonymization techniques in R. (3h)

LMU OSC logo
OSC Tutorial

Synthetic Data

Synthetic data creation in R to balance utility and privacy when sharing data. (3h)

2.3.3 GDPR Compliance

Research at LMU Munich must comply with EU data protection regulations. The core principles: have a lawful basis for processing personal data (usually consent or legitimate research interest), use data only for stated purposes, collect only what you need, delete data when you no longer need it, and protect it against unauthorized access.

In practice: document your lawful basis, include data protection language in consent forms, use institutional storage rather than personal cloud services, restrict access to those who need it, and plan when and how you will delete data.

For data protection guidance, contact the LMU Data Protection Officer or the Research Data Management team of the University Library.

2.4 Quality Control

Quality control catches problems before they propagate into your analysis. The practices here ensure your data is trustworthy and your exclusions are well-founded.

Define criteria before looking at your data. This prevents unconscious bias in what you keep and exclude, and demonstrates that your decisions are principled rather than convenient (see 1.4. Study Design & Analysis Plan).

2.4.1 Validation

Validation checks whether your data meets specifications. For example, values can be verified to fall within valid ranges (e.g., age ≥ 0), required fields can be checked for missing entries, categorical variables can be restricted to allowed levels (e.g., “yes/no”), dates can be validated for proper format, and cross-variable constraints can be enforced (e.g., discharge date ≥ admission date). Run checks during collection to catch problems immediately, after collection for systematic review, and after any processing to verify transformations worked correctly.

  • Automate what you can. Check that data types are correct, values fall within expected ranges, required fields are populated, and formats are consistent. These checks should run automatically and flag problems for review.

  • Catch what automation misses with manual review. Sample your data and verify it against the source. Inspect outliers to determine whether they are errors or genuine extreme values. Look for suspicious patterns: survey responses that alternate predictably, reaction times that are impossibly fast.

LEARN MORE

LMU OSC logo
OSC Tutorial

Data Documentation & Validation

Create validation rules and automated checks for your research data. (1h)

2.4.2 Cleaning

Data cleaning handles errors, inconsistencies, and missing values. The cardinal rule: never modify your raw data. All cleaning happens on copies, ideally by scripts and not through manual changes.

  • Correct unambiguous errors. Clear typos, obvious data entry mistakes. For ambiguous cases, flag them for review rather than making assumptions. Document your reasoning for every judgment call.

  • Handle missing data consistently. Decide on a coding scheme (NA, -999, blank) and apply it uniformly. When you know why data is missing, record that information. It may matter for analysis.

  • Investigate outliers before acting. An extreme value might be an error, or it might be genuine. Understand the cause before deciding whether to remove, transform, or retain it.

  • Write cleaning as a script. A script documents exactly what you did and lets you reproduce it. Keep a decision log for choices that cannot be automated.

2.4.3 Exclusions

Exclusion criteria specify which data points will be removed from analysis and why. Define these before you see your results (see 1.4. Study Design & Analysis Plan).

  • Review common exclusion criteria. Technical failures (equipment malfunction, incomplete recording), protocol violations (wrong procedure followed, participant did not comply), quality thresholds (too much missing data, failed attention checks), and participant criteria (did not meet stated inclusion criteria).

  • Document everything. Record criteria before analysis begins. Report how many data points were excluded for each criterion in a log with the date and reviewer’s identity. Plan sensitivity analyses comparing results with and without exclusions to show your findings are robust.

Collect & Manage Checklist

To complete before conducting a data validation meeting with members of the project. Not all items are relevant for all fields of research or study types.

Throughout Data Collection

Before Moving to Analysis

Download checklist

3 Analyze & Collaborate

4. Preserve & Share3. Analyze & Collaborate2. Collect & Manage1. Plan & Design

Create reproducible analyses and collaborate effectively

Readable Code Version Control Reproducible Computational Environment Reproducible Manuscript
Checkpoints: Repository & Code Peer-Review + Result Presentation

3.1 Data Processing & Analysis

Data processing and analysis should be reproducible – independent of which software, programming language, or operating system you use. This is best achieved by choosing automated script-based workflows (over manual point-and-click procedures), and supported by adequate documentation and shared code that allows others to regenerate results. Because analysis scripts are run repeatedly, due to iterative development through corrections and refinements, automation is essential for both reproducibility and efficiency. This section outlines a recommended workflow for R users.

3.1.1 Programming

  • Create a self-contained project folder. Include data, code, documentation, and outputs in a single structured environment, ensuring the project remains understandable, reproducible, and portable across systems and collaborators. If you use the free and open source software RStudio to manage your R project, your project directory (or folder) should contain a .Rproj file (see R tutorial). Use relative paths (i.e. “./subfolder”, where . represents the root of your .Rproj directory) or the library here, so the project stays portable to another environment
  • Use a standard folder structure. Your code repository should include a standard folder structure that make sense for your type of research, ideally shared across your team members. You can for instance use our research project template.
  • Stop clicking, start coding. Automatize all possible steps, including data acquisition (see 2.1. Data Collection), data processing and transformation, data analyses, data visualization, and results reporting (see 3.2. Reporting Results)
  • Structure, comment, and standardize your scripts. R scripts themselves should follow current standards to increase their readability (see Readable Code Lecture). Use meaningful names for variables, functions, and scripts. Add comments to your code explaining why you made a decision, any known limitations to your code, and citations of methods. Do not include sensitive information such as credentials or name of excluded patient as comments in your code!
    • Define your own functions rather than copy pasting pieces of code which makes it hard to maintain error-free. Functions are ‘self-contained’ sets of commands that accomplish a specific task. They usually ‘take in’ data or parameter values (these inputs are called ‘function arguments’), process them, and ‘return’ a result. See our R tutorial and data simulation tutorial for examples.
    • Set seeds for random processes to enable exact replication. A seed is a number used to initialize a pseudorandom number generator algorithm. It serves as the starting point for a sequence of numbers that appear random but are actually produced by a deterministic, fixed algorithm. See e.g. our data simulation tutorial for examples.
    • Follow accessibility standards when generating outputs (e.g. use colorblind-friendly color scheme for figures)
    • Follow a style guide to increase readability. Use automated styling tools (e.g. styler, lintr, and Air for R).
  • Use LRZ Compute Cloud for data-intensive analyses. LRZ Supercomputing provide virtual machines, high-performance computing, and storage to researchers of LMU Munich.

LEARN MORE

LMU OSC logo
R programming language logo
OSC Tutorial

R Tutorial

Process your data reproducibly in R. (3h)

LMU OSC logo
OSC Tutorial

Data simulation in R

Set seeds and define functions to generate random data (2h)

LMU OSC logo
OSC Lecture

Readable code lecture

Tips to write clear, understandable, and maintainable code. (1h)

TOOLS & RESOURCES

LMU OSC logo
GitHub logo
OSC Tool

Research Project Template

Project folder separating raw & processed data, code, and outputs.

LRZ logo
Supported at LMU

LRZ Supercomputing

Virtual machines and HPC for LMU Munich

3.1.2 Version Control

Version control tracks changes to files over time. You can see what changed, when, and why. You can revert to previous versions. Collaborators can work without overwriting each other.

In a version controlled workflow, you back up your local Git repositories on the cloud-based platforms GitHub or LRZ GitLab and share access to the online version of your repositories with your collaborators.

  • Learn to create git branches to collaborate on the same piece of code in a unique repository. i.e. temporary copies where you can work without breaking the original source code, which you later merge back to the main branch (see Advanced Git tutorial).

  • Maintain efficient communication to coordinate collaborative work. GitHub or GitLab facilitate collaboration through “issues” (a precise description of something to fix), “discussions” (asynchronous thinking through figuring out how to resolve a problem), and can still very well resolve “conflicts” (i.e. collaborators wanting to merge changes on the exact same line of script). To complement this, LMU Munich offers LMU chat (Matrix), a secured open source chat service with all LMU members on which you can also invite external collaborators.

NoteDifference between Git, GitHub and GitLab
  • Git is a version control system that tracks changes in text files (e.g. CSV, plain text, R, Python). The Git software and your Git repositories should be, respectively, installed and located in your local environment (i.e. on your computer, not on a drive, see Git tutorial).
  • GitHub is the most popular, free but proprietary and US-based cloud-based platform for software development with Git, providing collaboration features like pull requests and issues (see GitHub tutorial). You should not have any sensitive information on GitHub even in a private repository.
  • LRZ GitLab is a cloud-based hosting platform that works exactly the same as GitHub but is free and open source and is installed on the LRZ servers for LMU Munich and can therefore be considered secure when the repository is private.

While your LRZ GitLab account is associated with your LMU Munich affiliation, your GitHub account can be associated with your private email, be included in your CV, and be used for public sharing of your data and code (see 4. Preserve & Share).

In a version controlled workflow, you back up your local Git repositories on either GitHub or LRZ GitLab through a secure SSH connection (see GitHub tutorial) and share access to your repositories with your collaborators through the cloud-based platform GitHub or LRZ GitLab.

Important

If you work with sensitive data, you must not include the raw or processed data in the version-controlled repository that will end up being shared publicly.

Instead, explicitly exclude the data directory using the .gitignore file from the start, or, at the time of sharing, create a new local repository that contains all project files except the data.

Importantly, if data are removed from an existing repository, they may still remain accessible in the repository’s history, since previous states of the project can be restored. If sensitive data are accidentally committed and pushed, it is possible to rewrite the repository history to remove them retrospectively. However, this process is complex and error-prone, so it is best avoided by ensuring that sensitive data are excluded from version control from the outset.

TipTips for research groups:

Create a LRZ GitLab “organization” for the team. This allows repositories, permissions, and project resources to be managed centrally rather than under individual accounts. This ensures continuity when team members leave, as ownership can be transferred to e.g. the PI and other administrators of the organization.

LEARN MORE

LMU OSC logo
Git logo
OSC Tutorial

Git Tutorial

Use version control system Git from within RStudio. (2h)

LMU OSC logo
GitHub logo
OSC Tutorial

GitHub Tutorial

Collaborative coding with GitHub and RStudio (1h)

TOOLS & RESOURCES

LMU OSC logo
GitLab logo
Supported at LMU

LRZ Gitlab

Institutional Git hosting for LMU Munich.

LMU OSC logo
Matrix logo
Supported at LMU

LMU Chat (Matrix)

Institutional open source chat service for LMU Munich.

3.1.3 Computational Environment Management

Manage your computational environment by explicitly recording the software, package versions, and dependencies required for your analyses, ensuring results can be reproduced across systems and over time. Tools such as packages managers (e.g. Renv for R packages, Conda for Python packages) or broader containers (e.g. Docker or Binder) help stabilize workflows and prevent inconsistencies caused by packages or software updates.

For a R project repository:

  • Activate Renv to keep track of all packages versions (see our renv tutorial). This way, you or someone else can reproduce your results on another computer or at a later time using the same R packages versions.

Before publishing your project (see 4. Preserve & share):

  • Record your dependencies in your README file for possible reconstruction with repo2docker or binder (see Code Publishing tutorial).

LEARN MORE

LMU OSC logo
R package 'renv' icon
OSC Tutorial

Renv Tutorial

Manage the version of your R packages (1h)

LMU OSC logo
Zenodo icon
OSC Tutorial

Code Publishing Tutorial

Publish your reproducible project on Zenodo (2h)

TOOLS & RESOURCES

Binder logo

Binder

Create online executable environments.

Docker logo

Docker

Containerization for full reproducibility.

repo2docker logo

repo2docker

Build and run user environment containers.

3.1.4 Documentation

As with all documentation, your project repository’s documentation should be written early - initially for your near-future self to support efficient re-engagement after interruptions, then revised for internal team review, and ultimately expanded and refined for public sharing (see 4.2. Open Source Code).

  • Create a README (e.g. a .md or .txt file) early and update it as you go. Your README is the entry point to your project. A good README answers the essential questions: who created the script, what it contains, how they relate to other scripts and in which order scripts should be run, what the dependencies of the project are, how to obtain/access the input data, whether the code can be reused.
  • Annotate your code explaining why you made a decision. All parameter values used as input for a function, or other decisions, should be justified minimally as comments in your code to later be included in your manuscript. Do not include sensitive information such as credentials or name of excluded patient as comments in your code!
  • Update your data dictionary and README files. Your data documentation should be updated to include all data exclusion, change in range of possible values, etc. (see 2.2.4. Documentation).

Example README to allow team members to review your code:

# Analysis of Treatment Effects

## Requirements
- R version 4.3+
- Packages listed in renv.lock

## Running the Analysis
1. Install dependencies: `renv::restore()`
2. Run scripts in order: 01_preprocessing.R, 02_analysis.R

LEARN MORE

LMU OSC logo
OSC Lecture

Readable code lecture

Tips to write clear, understandable, and maintainable code.(1h)

LMU OSC logo
OSC Tutorial

Data Documentation & Validation

Create validation rules and automated checks for your research data. (1h)

LMU OSC logo
Zenodo icon
OSC Tutorial

Code Publishing Tutorial

Publish your reproducible project on Zenodo (2h)

3.2 Reporting Results

Your results should be computationally reproducible to ensure that the analysis can be independently verified. They should also be reported comprehensively and in line with reporting guidelines to facilitate statistical checks and future meta-analyses.

3.2.1 Results reproducibility

A practical way to make results and analysis decisions transparent and traceable is to use literate programming, which combines narrative text explaining the logic of the analysis with executable code that performs data processing and statistical procedures. When the document is compiled, outputs such as tables, figures, and references are generated directly from the code and updated automatically whenever the code changes, ensuring that the reported methods and results remain consistent with the analysis. This eliminates repeated copy-and-paste and reduces the risk of uncertainty about which version of, for example, a figure is actually included in the manuscript.

For users of R, Python, and Julia, Quarto has become the standard tool for creating reproducible reports, superseding R Markdown and enabling rendering to formats such as HTML, PDF, and Microsoft Word (see our Quarto tutorial).

Reproducible Analysis Report: The Minimum Standard for Computational Reproducibility
  • Create an analysis report with Quarto. To make your results reproducible, you should provide the analysis code that creates the results of the manuscript with explanations for each of the steps. Quarto is an ideal tool for this (see Quarto tutorial). Your analysis report should contain your research question, data loading instructions, preprocessing steps, statistical procedures, tables and figures, session info and software versions (in the report itself or in your README).

This will facilitate your own revisions, code-review by your team, reproducibility checks that some journals conduct as part of peer-review, and verification of results reproducibility by other researchers. The code repository can be published with simulated, synthetic, or real data, depending on the project, together with the respective article (see 4.2. Open Source Code).

Creating a reproducible analysis report in Quarto currently represents state-of-the-art practice for demonstrating the computational reproducibility of a study, and we recommend adopting this approach as a minimum standard.

Quarto can also be used to generate additional reporting outputs, such as manuscripts, presentations, and websites, in a fully reproducible manner. However, adopting a fully reproducible workflow for all outputs will require additional coordination and time, particularly when working with collaborators who use different tools or workflows.

Reproducible Manuscript: The Ultimate Reproducible Report

If you do not want to break the reproducibility chain by copy and pasting new results, tables, and figures into e.g. a Word document, you can write your entire manuscript in Quarto.

  • Include bibliographic references in your Quarto document. You can directly source .bib files into your Quarto document or use the open source reference management software Zotero integrated in RStudio to cite articles and have them automatically formatted in any journal standards (see Quarto tutorial and Zotero tutorial).

  • Use a template document for major formatting aspects. You can use e.g. a Word template to create your output with e.g. specific header and legend formatting. Some journals offer Quarto templates for further formatting requirements (see list of Quarto extensions). Some journals offer LaTeX or Word templates; you can adapt these by modifying Quarto’s YAML header and markdown content to create your own Quarto template.

  • Adapt your workflow to your collaborators. Ideally, all your team mates have also adopted the use of Git and GitHub or LRZ GitLab and the use of issues and branches to work collaboratively (see 3.1.2. Version Control). A compromise with collaborators who do not use Git is to render your draft manuscript to e.g. Word and share it through cloud-based collaborative document editing e.g. LRZ Sync & Share or Google Docs with collaborators using the suggestion / track changes mode. Then, the lead author transcribes all edits manually, or with the help of e.g. R packages such as trackdown or officer, and addresses all comments in their quarto version before re-rendering for a second round of revisions, potentially with additional analyses.

Reproducible Presentations & Websites: Further Reproducible Outputs for Outreach

Using the same Quarto environment, you can also:

  • Convert your analyses report into a reproducible presentation. To get feedback on your analyses, you can output a report into a slide presentation with presenter notes. Slides remain dynamically linked to the underlying analyses. Any modification to the data or code automatically propagates to tables, visualizations, and reported results, eliminating inconsistencies between presentation and computation.

  • Create reproducible website to support your research or teaching. Quarto also enables the creation of reproducible and interactive websites by linking content directly to data, code, and analyses within a unified publishing workflow. Pages, figures, tables, and teaching materials update automatically when the underlying sources change, ensuring consistency across outputs. This approach simplifies maintenance, and allows websites to evolve as living, version-controlled resources rather than static collections of files. Examples of websites, interactive reports and other formats can be explored on the Quarto Gallery. This website and our self-paced tutorials are all Quarto websites.

LEARN MORE

LMU OSC logo
Quarto logo
OSC Tutorial

Quarto Tutorial

Reproducible documents combining code and narrative (2h)

LMU OSC logo
Zotero logo
OSC Tutorial

Introduction to Zotero

Use a reference manager that integrates to RStudio. (1h)

LMU OSC logo
Git logo
OSC Tutorial

Git Tutorial

Use version control system Git from within RStudio. (2h)

LMU OSC logo
GitHub logo
OSC Tutorial

GitHub Tutorial

Collaborative coding with GitHub and RStudio (1h)

Git logo

Advanced Git Tutorial

Use Git branches and further collaborative features (2h)

TOOLS & RESOURCES

LMU OSC logo
GitLab logo
Supported at LMU

LRZ Gitlab

Institutional Git hosting for LMU Munich.

LRZ logo
Supported at LMU

LRZ Sync & Share

Cloud storage service for LMU Munich

LMU OSC logo
Matrix logo
Supported at LMU

LMU Chat (Matrix)

Institutional open source chat service for LMU Munich.

3.2.2 Reporting guidelines

Transparent and complete reporting is essential for the interpretation, verification, and reuse of scientific results. To support this, many disciplines have developed reporting guidelines that specify the minimum information that should be included when describing study design, data collection, analysis, and results. These guidelines help ensure that studies can be critically evaluated, replicated, and included in evidence syntheses such as systematic reviews and meta-analyses.

A central resource for such guidelines is the EQUATOR Network, an international initiative collecting and promoting reporting standards across many study types and disciplines.

3.2.2.1 Typical items that should be reported include:

  • Study design and setting – the type of study (e.g., randomized trial, cohort study), where and when it was conducted, and relevant contextual factors.
  • Participants or data sources – eligibility criteria, recruitment or sampling procedures, sample size, and reasons for exclusions or dropouts.
  • Variables and measurements – definitions of exposures, outcomes, predictors, and covariates, including how and when they were measured.
  • Sample size justification – power calculations or other reasoning behind the chosen sample size.
  • Statistical methods – the models used, assumptions checked, handling of missing data, variable transformations, and any sensitivity or robustness analyses.
  • Data preprocessing and analysis workflow – steps such as cleaning, filtering, or derived variables that affect the final dataset used for analysis.
  • Results with appropriate uncertainty – effect estimates, confidence intervals, p‑values, and clear descriptions of the comparisons performed.
  • Participant flow and descriptive statistics – numbers of observations at each stage and summary statistics of the analyzed sample.
  • Limitations and potential sources of bias – issues such as confounding, measurement error, or selection bias.
  • Data, code, and materials availability – where readers can access datasets, analysis scripts, or supplementary materials to reproduce the analysis.

A lot of these elements are those of a preregistration (see 1.4.1. Pre-analysis planning). Describe them in your methods section, adding they were planned a priori in your preregistration and citing its DOI, and separate results into e.g. preregistered confirmatory analyses and non-preregistered exploratory analyses.

Some of these elements will be refined depending on your study type. Widely used examples of guidelines include CONSORT for randomized controlled trials, ARRIVE for animal studies, and PRISMA for systematic reviews.

TOOLS & RESOURCES

EQUATOR logo

EQUATOR Network

Comprehensive database of reporting guidelines.

Analyze & Collaborate Checklist

This assumes a reproducible workflow using R. Not all items are relevant for all fields of research or study types.

1. During data analyses: create and maintain a repository

Repository Setup

Repository Content

R Scripts

2. Before presenting results to the research group: conduct internal code peer review

3. Before sharing outside the group

Download checklist

4 Preserve & Share

4. Preserve & Share3. Analyze & Collaborate2. Collect & Manage1. Plan & Design

Make your research outputs accessible, reusable, and citable

Data Sharing Code Publishing Open Access Persistent Identifiers
Checkpoint: Manuscript submission with DOIs for preregistration, data, and code

4.1 FAIR Data Sharing

Sharing your data allows others to verify your findings, build on your work, and increase the impact of your research. But sharing does not always mean making everything publicly available. What you can share depends on the consent you obtained, the sensitivity of your data, and your ethics approval.

In 1. Plan & Design, you wrote a Data Management Plan, in 2. Collect & Manage, you created documentation for yourself, and in 3. Analyze & Collaborate you updated that documentation for your collaborators. Now you prepare that documentation for people who have no prior knowledge of your project and reassess you plan:

  • Make a final decision on how openly your data will be shared.
  • Select an appropriate repository for publishing your data type.
  • Apply a license that tells others how they may reuse your work.

These steps are straightforward if you followed previous recommendations: by this stage, you should already have a README, a data dictionary (codebook), organized files, and clear metadata.

4.1.1 Open vs Restricted

Not all data can or should be shared openly. Your sharing options depend on what consent participants gave and what your ethics approval permits (see 1.2.3. Ethics).

  • Open data can be downloaded by anyone without restrictions. This maximizes reuse but is only appropriate when data contains no personal information or has been fully anonymized.

  • Restricted access means the metadata is available on a professional repository and clear instructions of the conditions of access to the data are provided. Requesting the data often requires describing the purpose of reuse (sometimes in the form of a preregistration, see 1.4.1. Pre-analysis planning), signing a Data Use Agreement, and agreeing to further conditions. This can be organized automatically through the repository or through contacting the authors. This balances reuse with protection.

  • Metadata only describes what data exists without sharing the data itself. Others can discover your work and contact you to discuss access. This is appropriate when data cannot be shared due to legal or ethical constraints. Making your data Findable by sharing your metadata on a professional repository is the first step to making your data FAIR, as emphasized by the LMU Guidelines for Safeguarding Good Scientific Practice (see also 1.2. Legal Requirements)

When deciding whether data should be shared, consider the following:

  • Ensure that data sharing is compatible with the informed consent provided by participants.
  • Check what your ethics approval allows.
  • Consider whether the data can be anonymized without losing scientific value (see 2.3.2. Anonymization).
  • If you work with sensitive data, you may consult the University’s data protection officer.
  • If your research outputs could have dual-use implications (e.g. that could also be applied for military or malicious purposes), consult the relevant regulations (e.g. European Commission’s policy (Dual-Use Regulation 2021/821)) and contact the University’s Export Control service.
  • If your research may lead to patents or commercialization, contact the the University’s IP Management team before sharing data. Early consultation helps ensure that intellectual property rights are not compromised.

LEARN MORE

LMU OSC logo
OSC Lecture

Why share data openly?

An introduction to the what, why, and how to make data open (30 min)

LMU OSC logo
OSC Lecture

Maintaining Privacy with Open Data

How to make data open without revealing sensitive information (1h)

LMU OSC logo
OSC Tutorial

Data Anonymization

Implement data anonymization techniques in R. (3h)

TOOLS & RESOURCES

LMU OSC logo

LMU Guidelines for Safeguarding Good Scientific Practice

Implementation of the German Research Foundation's (DFG) Code of Conduct

European Union flag

EU Dual-Use Regulation 2021/821

List of regulated research outputs that could lead to military or malicious purposes.

4.1.2 Preparing Your Data

During the course of data collection and analyses, you created a README and data dictionary for yourself and collaborators. Before sharing publicly, review them from the perspective of someone who knows nothing about your project.

  • Expand your README to include the research context (who created the data, what it contains, when and where it was collected, why it was generated, how it was produced), how to cite the dataset, and any access conditions.

  • Review your data dictionary to ensure every variable is fully described. What was obvious to you during analysis may need explanation for others.

  • Confirm you follow community standards. Many fields have established formats for sharing data (e.g., BIDS for neuroimaging, see FAIRsharing and RDMkit to review discipline specific organizational and metadata standards). Using these makes your data immediately usable with existing tools (see also 2.2.5. Standards).

  • Verify anonymization. Check that no personal information remains in the data files, metadata, or filenames.

LEARN MORE

LMU OSC logo
OSC Tutorial

FAIR Data Management

Manage data following FAIR principles. Covers READMEs, data dictionaries, file formats, and standards. (2h)

LMU OSC logo
OSC Tutorial

Data Documentation & Validation

Create data dictionaries and READMEs for your data. (1h)

LMU OSC logo
OSC Tutorial

Data Anonymization

Implement data anonymization techniques in R. (3h)

TOOLS & RESOURCES

FAIRsharing logo

FAIRsharing

Search by discipline to find metadata standards, reporting guidelines, and data policies for your field.

RDMkit logo

RDMkit

Domain-specific metadata standards

4.1.3 Where to Deposit

Choose a repository that suits your data type, your field’s expectations, and your access requirements.

  • Discipline-specific repositories are often the best choice. They use metadata standards your community expects, making your data findable by researchers in your field. Search re3data to find repositories for your domain.

  • General-purpose repositories like Zenodo or OSF accept any data type. They provide DOIs and long-term preservation, but may lack the specialized metadata fields of discipline-specific options.

  • Institutional repositories may be required by your funder. Choosing the LMU repository Open Data LMU also ensures you can get the support of the Research Data Management team of the University Library.

Whichever you choose, ensure the repository provides a DOI (Digital Object Identifier) so your data can be cited and tracked.

TOOLS & RESOURCES

re3data icon

re3data

Registry of research data repositories.

OSF icon

Open Science Framework

General-purpose repository for data, materials, reports.

Zenodo icon

Zenodo

General-purpose repository for data, software, reports.

LMU logo
Supported at LMU

Open Data LMU

Institutional data repository.

4.1.4 Data Licenses

A license tells others what they can do with your data. Licensing your data consists in adding a file called LICENSE.txt next to your data, that contains the appropriate legal text. Without one, or equivalent statements, others cannot legally redistribute your research outputs, or reuse them in work they publish themselves, even if they are publicly available.

Common open licenses for data are:

  • CC0 (Public Domain) places no restrictions. Anyone can use, modify, and redistribute without attribution. Recommended for factual data where maximum reuse is the goal.
  • CC BY (Attribution) requires users to credit you. A good default when you want recognition while enabling broad reuse.
  • CC BY-NC (Non-Commercial) adds a restriction against commercial use. Consider whether this limitation actually serves your goals.
  • A Scientific Use License (Public Use) only allows re-use for scientific purposes (but with all freedoms within that scope). The Leibniz Institute for Psychology (ZPID) developed such a standard license together with legal experts.

Learn more about open licenses for data and code in our code publishing tutorial.

For restricted-access data, a Data Use Agreement specifies conditions beyond a standard license: approved purposes, security requirements, publication terms, and data destruction timelines.

LEARN MORE

LMU OSC logo
OSC Tutorial

Choose a License

License decision flowchart for data and code.

4.1.5 Data Use Agreements

in construction

  • Description of what it is, when it is needed, what the process to create one is (Team SOP for handling access requests (who receives them, how decisions are made, what must be documented, and response timelines))
  • LMU contact for legal department
  • Examples DUA

4.2 Open Source Code

Making code publicly available demonstrates the reproducibility of your results and enables others to understand, verify, and build upon your analytical methods.

So far, your code was either backed up on the secured LRZ Gitlab (e.g. if your data and/or code are sensitive), or on GitHub (see 3. Analyze & Collaborate). Before the submission of a manuscript to a journal and/or upon the acceptance of a manuscript, there are small additional steps that need to be done to publish your code.

  • verify the structure of your repository, the readability of your scripts, the completeness of the documentation (see Analyze & Collaborate Checklist).
  • make a clean version public, e.g. on GitHub
  • add a license
  • get a DOI, e.g. through Zenodo

If you work with sensitive data that cannot be anonymized and shared:

  • generate a simulated random dataset to allow for the published code to run (which you may have already done if you simulated data in order to prepare a preregistration, see 1.4. Study Design & Analysis Plan), or
  • create a synthetic dataset with the same properties as the original dataset to allow others to re-derive an approximation of the original results and conduct further exploratory analyses.

4.2.1 Preparing Your Code Repository

During the course of data analyses, you created scripts and documentation such as README and data dictionary for yourself and collaborators. Before sharing publicly, review them from the perspective of someone who knows nothing about your project.

  • Expand your README to include a description, involved data, computational requirements and dependencies (i.e. what software, packages, and their version, need to be installed to run the analyses), list of results.
  • Document your code from an external perspective, using literate programming (e.g. Quarto) or comments.
  • Double check that no sensitive information remains in your repository (e.g. code comments, history of sensitive data)

See our code publishing tutorial for more information on how to prepare your code repository for sharing.

Important

If you work with sensitive data, you must not include the raw or processed data in the version-controlled repository to be shared.

Instead, explicitly exclude the data directory using the .gitignore file from the start. An easy solution (which, however, discards the version history for all files), is to create a new local repository that contains all project files except the data and only push that to the public repository.

Importantly, if data are removed from an existing repository, they may still remain accessible in the repository’s history, since previous states of the project can be restored. If sensitive data are accidentally committed and pushed, it is possible to rewrite the repository history to remove them retrospectively. However, this process is complex and error-prone, so it is best avoided by ensuring that sensitive data are excluded from version control from the outset.

LEARN MORE

LMU OSC logo
OSC Tutorial

Code Publishing

Add all elements of a reproducible project to your repository.

4.2.2 Real, Simulated, or Synthetic Data

Sharing your code allows for other researchers to clearly see which analytic methods were applied to the data. Ideally, they should also be able to rerun the code and verify the reproducibility of the reported results.

You can provide the data required to run the code in several ways:

  • The real (anonymized) data. This can be done either (1) by including the dataset in the repository so that the code can access it locally and run directly, or (2) by configuring the code to retrieve the data from an external source, such as a database, API, or external data repository. A practical workflow is to include “small”/one-shot datasets directly into a combined code/data-project (e.g. on GitHub), and to store “large”/reusable datasets (which deserve their own DOI) separately in a repository specialized for research data.
  • Simulated random data. This option mainly demonstrates that the code runs without errors. However, the original results cannot be verified and further analyses are not meaningful. See our data simulation tutorial.
  • Synthetic data that mimic key properties of the real data. When privacy-sensitive data cannot be shared, synthetic datasets can provide a useful alternative. Compared with purely random simulated data, synthetic data can resemble the structure and characteristics of the original dataset, allowing other researchers to rerun the analyses and assess whether the main results can be reproduced. In addition, openly shared synthetic data enable exploratory analyses that may generate new hypotheses, and in some cases analyses conducted on synthetic data can approximate results obtained from the real data.

LEARN MORE

LMU OSC logo
OSC Tutorial

FAIR Data Management

Manage data following FAIR principles. Covers READMEs, data dictionaries, file formats, and standards. (2h)

LMU OSC logo
OSC Tutorial

Data simulation in R

Easy data simulations in R. (2h)

LMU OSC logo
OSC Tutorial

Synthetic Data

Synthetic data creation in R to balance utility and privacy when sharing data. (3h)

4.2.3 Code Licenses

A license tells others what they can do with your code. Licensing your code consists in adding a file called LICENSE.txt next to your code, that contains the appropriate legal text. Without one, or equivalent statements, others cannot legally reuse your code, even if it is publicly available e.g. on GitHub. Common open licenses for code are:

  • CC0 (Public Domain) places no restrictions. Anyone can use, modify, and redistribute without attribution. Recommended for generic code where maximum reuse is the goal.
  • MIT is a simple license that requires users to credit you when they reuse, modify, and redistribute your code.
  • Apache 2.0 is a safer license (covering more legal cases) that requires users to credit you and state the changes you made to the code.

Learn more about open licenses for data and code in our code publishing tutorial or using the tool ‘Choose an open source license’.

LEARN MORE

LMU OSC logo
OSC Tutorial

Choose a License

License decision flowchart for data and code.

Choose an open source license logo

Choose an open source license

Tool comparing open software licenses.

4.2.4 Archiving & DOIs

To publish your code on Zenodo, we recommend to

  • Push your clean repository to GitHub (see GitHub tutorial).
  • Create an account on Zenodo.
  • Link your GitHub account to Zenodo. Navigate to your profile > account settings > link external accounts > GitHub > authorize Zenodo to access your GitHub account > a list of your GitHub repositories should appear. Enable the repository by toggling the switch next to it; refresh the page to check to see if the visual indicator that the repository is connected appears
  • Go to GitHub and create a release. Zenodo will automatically download a .zip-ball of each new release and register a DOI.
  • Add DOI badge to your README. After your first release, a DOI badge that you can include in GitHub README will appear next to your repository in your list of GitHub repositories on Zenodo. This allows others to easily find and cite your archived code
  • Create new release when needed. If you make a new release, it will create a new version on Zenodo, under the same overall DOI, with a ‘sub-DOI’ to identify a specific version.
TipTips for research groups:

Create a Zenodo “community” for the team. After team members connect their repositories from their personal GitHub account to Zenodo, and archive releases with a DOI, these records can be added to a shared Zenodo community. A community provides a central space where all software and other research outputs produced by the group can be collected and displayed, making them easier to find, cite, and showcase at the team level.

LEARN MORE

LMU OSC logo
OSC Tutorial

Code Publishing

Add all elements of a reproducible project to your repository.

TOOLS & RESOURCES

Zenodo icon

Zenodo

General-purpose repository for data, software, reports.

4.3 Open Materials

Any material needed to reproduce or replicate your study should also be shared openly unless there are dual-use, patent, or privacy concerns.

4.3.1 Digital materials

For instance, you can share your:

  • Share wet-lab protocols on e.g. protocols.io, see 2.1.1. Lab Protocols.
  • Share survey text, instructions, or scoring sheets on e.g. OSF

We recommend to use a creative Common license such as

  • CC-BY to allow reuse, modification and redistribution, while getting attribution; or one of its derivatives e.g.
  • CC-BY-SA (ShareAlike) under which credit must be given to the creator and adaptations must be shared under the same terms.

See the Creative Common license chooser to explore more.

TOOLS & RESOURCES

Protocols.io icon

Protocols.io

Share, discover, cite, and improve research protocols.

OSF logo

OSF

Project management platform including storage and DOI.

Creative Commons logo

Creative Commons licenses

Tool comparing licenses for diverse creative work.

4.3.2 Physical materials

Repositories for physical research materials (e.g., biological samples, chemicals, specimens, hardware, or other tangible resources) are usually called biobanks, material repositories, or research infrastructure collections and are domain specific or specialized to one kind of material (e.g. DNA).

For instance, Addgene is a nonprofit plasmid repository deposited by thousands of research labs around the world, while ARIADNE’s mission is to create a sustainable infrastructure for archaeological data sharing across institutions and nations.

Regardless of your discipline:

  • Use repositories that assign persistent identifiers for proper citation.
  • Provide detailed metadata, including provenance, storage conditions, and usage restrictions.
  • Check legal and ethical requirements (biosafety, export control, patient consent).
  • Document links between materials and associated datasets or publications.

4.4 Open Access Articles

Publishing open access articles (i.e. which are free to readers) increases the visibility, citation, and impact of research by removing paywall barriers, enabling scholars, practitioners, and policymakers worldwide to access and build upon the work without restriction. It also accelerates knowledge dissemination and promotes equity in scholarship by providing institutions and researchers - especially in low-resource settings - free and immediate access to scientific findings. Finally, tax-payer funded projects often have particular requirements for open access.

4.4.1 Pathways to Open Access

There are several pathways to make your articles free to read, e.g.:

  • Diamond Open Access: journals that provide immediate open access without charging authors any Article Processing Charges (APCs), as publication costs are covered by institutions, consortia, or public funding. Visit the Directory of Open Access Journals (DOAJ) to find them.
  • Gold Open Access: the publication of the final version of record as immediately open access on the publisher’s platform, against payment of an Article Processing Charges (APCs) by the authors.
  • Green Open Access: self-archiving of a preprint or postprint in an institutional or disciplinary repository, either immediately or after an embargo period imposed by the publisher, without cost to the author. Visit the Open Policy Finder to know whether and where your publishers allow you to publish preprint or postprints.

Depending on the contract negotiated by the University and each publisher (see 4.4.2. Contract with Publishers), the payment of APCs can also be requested by publishers regardless of the article’s accessibility status.

NoteDefinitions

preprint: article, book chapter, book, or other scholarly work that is deposited in a repository with unrestricted access (i.e. available to the public to view online, or download, without registration, payment, or approval) ahead of peer review. Equivalent terms used in some disciplines are ‘working paper’ and ‘unpublished manuscript’.

postprint: accepted version of an article, book chapter, book, or other scholarly work that is released by the author(s) on a repository with unrestricted access (i.e. available to the public to view online, or download, without registration, payment, or approval)

TOOLS & RESOURCES

DOAJ logo

Directory of Open Access Journals (DOAJ)

Index of trusted open access journals.

Jisc logo

Jisc' Open Policy Finder

Clear summary of journals' open access policies.

4.4.2 Contracts with Publishers

Until 2024, authors publishing in legacy (subscription-based) journals were required to pay an Article Processing Charge (APC) to make their work immediately open access upon publication, the model commonly referred to as ‘gold open access’. To avoid these fees while still ensuring public accessibility, authors were encouraged to disseminate a preprint and/or postprint, or to deposit the publisher’s version in an institutional repository after the embargo period had expired. These procedures, which are free to the authors, are collectively known as ‘green open access’.

Since 2025, however, at LMU Munich and other universities that are part of the DEAL Konsortium in Germany, the agreements with Elsevier, Wiley, and Springer Nature, stipulate that the authors publishing in hybrid journals (i.e., journals offering both subscription-based and open access options) incur an Article Processing Charge (APC) regardless of whether the article is published closed or open access. Consequently, authors who elect to publish in these venues and pay the APC are generally advised to opt for immediate open access through the publisher. Although preprinting remains possible and may expedite dissemination, it no longer constitutes a cost-free alternative for hybrid journals operating under these uniform APC conditions.

The current contract conditions for LMU Munich can be accessed on the University Library Publication Fees Webpage (in german). To get support with Article Processing Charges and open access policies, contact the University Library Open Access team.

TOOLS & RESOURCES

LMU OSC logo
Supported at LMU

LMU University Library

Publisher contracts and publication fees.

4.4.3 Publishing Preprints and Retaining Rights

To increase the accessibility and impact of your article we recommend you to:

  • Check what your target journal’s open access policies are. Jisc’ Open Policy Finder provides clear information about whether you are allowed to archive various versions of manuscripts. For instance, your target journal may, or may not, allow you to deposit:

    • the version submitted to the journal (preprint)
    • the accepted manuscript version (postprint)
    • the published version
      • on a specific location (e.g. preprint repository, institutional repository, funders repository, author’s personal website)
      • after a specific amount of time (e.g. 6 or 12-month embargo)
      • under a specific license (e.g. CC-BY (recommended for maximum reuse), CC-BY-NC-ND)
  • Disclose the state of peer review on each preprint version. Use watermarks such as ‘Draft’ and ‘Revised Draft’ for your preprint, with a statement on the title page highlighting that the work has not been peer reviewed. For postprints, indicate that the version was accepted through peer review, and link to the DOI of the official journal version.

  • Get a DOI for each public version of your manuscript. When updating your preprint through the revision cycle, get a citable DOI (as provided by the repository).

  • Maintain as much rights to your own work as possible.

    • Prior to submitting your manuscript, deposit your manuscript’s conceptual figures on a repository (e.g. Zenodo, OSF), under an open license (e.g. CC.BY 4.0 allowing reuse and modification while requiring authors to get credited), and get a DOI for it to cite your own figure in your manuscript. This way, you (and possibly other researchers), can reuse them in future manuscripts without asking permission from the publisher, by indicating the DOI and license in the figure legend.
    • When given the choice, request your publisher to publish your manuscript under a CC.BY 4.0 license, so your right to distribution is included. It is possible that your publisher will request a CC.BY. NC (Non-Commercial) license which means you cannot deposit it on commercial platforms, but you would typically still be allowed to deposit it on preprint servers or institutional repositories (see Jisc’ Open Policy Finder for the policy of your specific journal)

Preprint flow

In practice, we recommend to follow the process of version releases illustrated in the figure, starting with a manuscript draft, should you wish to obtain feedback from your community, or preprinting the submitted version, should you want to e.g. accelerate dissemination of findings that are unlikely to change during revision, or postprinting the accepted version to have a free peer-reviewed version available to anyone. Note that submissions to some preprint servers are moderated (e.g. to prevent submission out of scope or generated by AI) which can take a few days.

To get support for Open Access LMU, the LMU Munich institutional repository for articles and books, contact the University Library Open Access team.

TOOLS & RESOURCES

Jisc logo

Jisc' Open Policy Finder

Clear summary of journals' open access policies.

LMU logo
Supported at LMU

Open Access LMU

Institutional publication repository.

OSF logo

OSF Preprints

Multidisciplinary preprint service.

bioRxiv

Biological sciences preprints.

medRxiv

Medical and health sciences preprints.

arXiv

Physics, mathematics, computer science preprints.

4.5 Attributions & Persistent Identifiers

All research outputs and their specific versions should be linked to their authors, organizations, and funders through persistent identifiers.

4.5.1 Authorship, Contributorship, and ORCID

Authorship

According to §14 of the LMU Guidelines for Safeguarding Good Scientific Practice, an author must:

  • make a genuine, verifiable contribution to the content of a scientific publication, data, or software, and not solely have a managerial or supervisory position in the project;
  • agree to the final version of the work to be published and bear joint responsibility for the publication unless explicitly stated otherwise.

Publishers can provide additional guidance on authorship decisions (often based on the recommendations of the International Committee of Medical Journal Editors.

We recommend to start discussions on authorship, roles, and responsibilities as early as possible. First discussions can for instance take place at the end of 1. Plan & Design, e.g. during the presentation of the study plan to the research group (see Plan & Design Checklist). As roles may shift in the course of a project, a review is warranted at the write-up stage.

Contributorship
  • Use the Contributor Role Taxonomy (CRediT) to increase transparency and accountability, and acknowledge what contribution to a research project each person has made: ‘Conceptualization’, ‘Data curation’, ‘Formal analysis’, ‘Funding acquisition’, ‘Investigation’, ‘Methodology’, ‘Project administration’, ‘Resources’, ‘Software’, ‘Supervision’, ‘Validation’, ‘Visualization’, ‘Writing—original draft’, and ‘Writing—review & editing’.
    • Use the Tenzing App to keep track of contributorships in large collaborations and export the information into various reusable formats.
    • Include the CRediT statement in the acknowledgement section of your manuscript if your publisher is not a formal CRediT adopter and does not otherwise prompt you upon submission of your article.
  • Use the Method Reporting with Initials for Transparency (MeRIT) when relevant. This consists in adding initials to the Methods section, identifying with greater transparency who did what when several team members contributed to e.g. data curation or software.
ORCID
  • Create an Open Researcher and Contributor ID (ORCID), a free, unique, persistent identifier for individuals, to distinguish yourself and claim credit for your work automatically, no matter how many people have your same (or similar) name.
  • Associate your ORCID to all your research outputs. The vast majority of publishers, funders, and data and code repositories prompt for authors and contributors’ ORCID. You can also sign in to many academic and institutional services with your ORCID (e.g. Zenodo, OSF, RDMO).

TOOLS & RESOURCES

LMU logo

LMU Guidelines for Safeguarding Good Scientific Practice

Implementation of the German Research Foundation's (DFG) Code of Conduct

CRediT logo

CRediT

Contributor Role Taxonomy.

ORCID logo

ORCID

Free, unique, persistent identifier for researchers.

4.5.2 Institution, Funders, and ROR

  • Disclose your institutional affiliations and funders as part of the metadata of all your research outputs. To be machine-readable and automatically connected to people’s profile or their organization, this must include the official name of the organization and ideally (one of) its persistent identifier.
  • Use your institution and funders’ ROR. The Research Organization Registry (ROR) is a global, community-led registry of open persistent identifiers for research and funding organizations.
Note

The official names of LMU Munich are Ludwig-Maximilians-Universität München in German, and LMU Munich in English. This university’s ROR ID is https://ror.org/05591te55.

LMU Munich is mostly funded by the German Research Foundation, whose ROR ID is https://ror.org/018mejw64

The LMU Open Science Center is one of the many child organizations of LMU Munich; its ROR ID is https://ror.org/029e6qe04. The funder of this handbook’s project is the Volkswagen Foundation, https://ror.org/03bsmfz84.

TOOLS & RESOURCES

ROR logo

Research Organization Registry (ROR)

Persistent identifiers for research and funding organizations

4.5.3 Connecting Your Work

  • Request a Digital Object Identifiers (DOI) for each of your research outputs (data, code, preprint, preregistration)
    • you automatically get a DOI from professional repository for outputs you make public
    • for a code repository you first want to keep private on e.g. Zenodo, you can request a DOI (by ticking a box) to reserve a DOI link and already cite it in your manuscript.
    • for a preregistration you first embargoed (i.e. kept private for a set period of time) on e.g. the OSF, you must first make it public to obtain a DOI to include in your manuscript.
  • Add persistent identifiers of connected outputs in their reciprocal metadata to connect all parts of a project.
    • include the DOI of your data, code, and preregistration in your manuscript and in the metadata field provided by your publisher if prompted.
    • add the DOI of your code repository in the metadata of your published data and reciprocally, either in the description text or in the metadata fields provided by repositories like Zenodo.
  • Add the persistent identifier of a published version of an output to all public versions of that output
    • add the DOI of the published article (provided by your publisher) in the metadata of your preprint.
    • add the DOI provided by e.g. Zenodo when publishing your code from GitHub, back into your GitHub README file.

Preserve & Share Checklist

Not all items are relevant for all fields of research or study types.

Before sharing resources

Upon publishing resources

Upon submitting an article

After acceptance of an article

Download checklist

Ludwig-Maximilians-Universität
LMU Open Science Center

Leopoldstr. 13
80802 München

Contact

  • Prof. Dr. Felix Schönbrodt (Managing Director)
  • Dr. Malika Ihle (Coordinator)
  • OSC team

Join Us

  • Subscribe to our announcement list
  • Become a member
  • Chat with us on Matrix

Imprint | Privacy Policy | Accessibility