← All resources

6 Files Labs Need to Share Analysis Securely Without Exposing Data

11 min read
6 Files Labs Need to Share Analysis Securely Without Exposing Data

6 Files Labs Need to Share Analysis Securely Without Exposing Data

Researcher reviewing restricted analysis files

Sharing analysis securely means publishing the complete computational pipeline, the data descriptors, code, environment snapshot, and provenance record, while applying the access level the data warrants, whether that is an open DOI or a controlled-access gateway. The first step is building a reproducible research compendium and deciding early whether your data can be open or must be restricted. Agentic analytics platforms like PlotStudio AI are increasingly part of that workflow, since they can plan, execute, and document the analysis as it happens rather than leaving provenance to be reconstructed later.


TL;DR:

  • Reproducible research compendiums should include raw and processed data descriptions, clean analysis scripts, environment files, and clear documentation to ensure full reproducibility.
  • Privacy assessments must be completed before packaging data, with automated scans for personally identifiable information, and sensitive data should be stored in controlled-access repositories.
  • Container images should be referenced by immutable digests, scanned for vulnerabilities, and archived with DOIs to guarantee long-term reproducibility and integrity of computational environments.
  • Publishing code and containers with persistent identifiers like DOIs, along with metadata and citation files, enhances discoverability and ensures proper attribution for reusing the work.
  • For sensitive data, open code and metadata should be published openly, while raw datasets are gated behind access requests, with clear vetting procedures and audit trails to prevent delays.

Plotstudio
Keep Research Analysis Reproducible
PlotStudio AI plans, runs, and documents multi-step analysis while keeping code, outputs, statistics, and visualizations available for review.

Table of Contents

Build the research compendium: a checklist of what to prepare

Before anything goes out to a collaborator or a journal, assemble the artifacts that let someone else rerun your work without guessing at your intentions. The goal is a self-contained package, not a folder of loose scripts.

  • Raw and processed data descriptors with variable definitions and provenance notes.
  • Clean, commented analysis scripts or notebooks, separated from exploratory work.
  • An environment file (environment.yml, renv.lock, or requirements.txt) pinned to exact versions.
  • A Dockerfile or Singularity recipe that rebuilds the computational environment.
  • A README explaining the folder structure, run order, and expected outputs.
  • A CITATION file so collaborators and reusers can credit the software correctly.

Structure the repository so inputs, code, and outputs live in separate, clearly named folders, then tag a release in Git before archiving it. That tagged release, with accompanying release notes describing what changed and why, becomes the version you cite. Journals increasingly expect this level of packaging: top venues now ask authors to share the full computational pipeline, code, data, environment, and documentation. They prefer executable capsules that referees can run directly, according to commentary in Nature Machine Intelligence.

Finish with two practical steps:

  1. Write a smoke test, a short script that runs the pipeline on a small sample and checks that outputs match expected hashes or figures.
  2. Record the exact command used to launch the full analysis, so a reviewer can reproduce it without reverse-engineering your notebook.

Frameworks like the CURE-FAIR “10 things” guidance cover this same ground in more depth if you want a curation checklist to follow line by line.

Protect participant privacy before anything leaves your machine

Privacy review happens before packaging, not after. Treat it as a gate, not an afterthought bolted onto the README.

  • Run a re-identification risk assessment on any dataset involving human subjects, even if it looks anonymized.
  • Document consent terms and legal restrictions that limit redistribution, and keep that documentation with the data descriptor.
  • Scan code and commit history for embedded PII (API keys, participant identifiers, email addresses) using automated repository scans, a practice recommended in UK Biobank’s guidance on code repositories.
  • Redact or aggregate sensitive fields rather than relying on dropping a single identifier column.
  • Write a data availability statement that names what is open, what is restricted, and exactly how to request access.

De-identification alone is rarely sufficient for sensitive human-subjects data. Controlled-access repositories paired with a documented vetting process offer stronger protection against re-identification than relying on removing obvious identifiers, per NIH guidance on data access and privacy.

Pro Tip: Run your PII scan on the full commit history, not just the latest snapshot. Deleted files and old commits often still contain what you meant to remove.

Illustration of scanning repository history

Package the execution environment: containers and image digests

Package managers like renv or pip capture library versions, but they miss operating-system libraries that can silently change numerical results years later. Container images solve that gap, which is why The Turing Way recommends combining both: package managers for fast iteration, containers for long-term portability.

Containerization is now considered a cornerstone of reproducible computational research, but it has to be paired with security discipline rather than treated as a packaging convenience, according to nine tips for container use in computational biology published in PLOS Computational Biology.

  • Reference images by immutable digest rather than a floating tag like “latest,” since tags can be overwritten while digests cannot, per Docker’s own documentation on image digests.
  • Use dual tagging: a “living” tag for active development and a digest-pinned, DOI-archived build for the version tied to publication.
  • Scan images for known vulnerabilities and run them with least-privilege or rootless settings before sharing.

Immutable image digests, archived in a DOI-capable repository, give the strongest guarantee that a container pulled five years from now matches the one you ran today, a practice described in Docker’s image digest documentation.

Publish and archive: repositories, PIDs, and software citation

A bare GitHub link is not a citable artifact. Repositories change, branches get deleted, and accounts get renamed, so a persistent identifier matters as much as the code itself.

  • Deposit code and container images in a repository that mints a DOI, such as Zenodo, Figshare, or Code Ocean, producing a permanent landing page.
  • Attach machine-readable metadata (author, version, license, dependencies) alongside the deposit so indexers and reusers can parse it automatically.
  • Include a CITATION file and a code availability statement that matches the specific wording your target journal requires.

Publishing code with a DOI, a CITATION file, and a clear availability statement improves discoverability and gives you proper credit when others reuse the work, a point echoed in Springer Nature’s guidance on sharing research code. The FAIR4RS principles for research software extend this logic further, calling for persistent identifiers and accessible metadata as baseline requirements rather than optional extras. A landing page also survives repository migrations in a way a raw repository URL never does.

Sharing restricted data: controlled-access workflows that still work

Not every dataset can be open, and pretending otherwise invites compliance problems—this is why understanding security and confidentiality practices is essential. The practical answer is to separate what can be public from what must be gated.

  1. Classify the dataset against consent terms, legal constraints, and sensitivity before deciding on an access tier, following the criteria in NIH’s guidance on designating controlled-access data.
  2. Publish code, metadata, and a data descriptor openly while gating only the raw sensitive data behind a controlled-access identifier.
  3. Build a lightweight access request template and a vetting checklist so approvals do not stall on ad hoc email threads.
  4. Log every grant and denial, including the reviewer and date, so the process is auditable later.

Controlled-access metadata should travel with the public artifacts, so a reader encountering your open code immediately understands how to request the gated portion and under what terms, per the same NIH DMS policy.

Pro Tip: Put the access request template directly in the repository README. Reviewers waiting on a buried email thread are the most common reason controlled-access requests stall for weeks.

Verification: provenance and smoke tests that let reviewers trust the result

A reviewer should be able to trace every output back to the exact input and command that produced it. Tools like DataLad can link containers directly to datasets and record the provenance of a run, so a collaborator on a different machine can execute the recorded command and get the same result, as described in the DataLad handbook.

  • Record commands, inputs, and outputs together rather than relying on memory or scattered notes.
  • Include a smoke test with expected hashes or figure checksums so a failed rerun is caught immediately, not discovered months later.
  • Document the usual failure points: dependency version drift, unset random seeds, and hardcoded local file paths.

A short troubleshooting note in the README, listing these three failure modes and how you resolved them, saves reviewers hours and saves you the follow-up emails.

How PlotStudio AI supports secure, reproducible sharing

The platform applies agentic analytics to this workflow: specialized agents plan an analysis, write and execute real code, inspect intermediate results, and assemble the findings into a package built for review rather than a one-off answer. Analyses can run locally, which matters directly for privacy-sensitive projects that should never touch a conventional cloud analytics platform.

  • A Plan Mode feature lets researchers review and adjust the proposed methodology, assumptions, and steps before any code runs.
  • Local execution keeps sensitive data on the user’s machine throughout the analysis.
  • Completed work exports to notebooks and PDF reports, preserving the code, statistics, and charts for review.
  • Domain-specific Skills encode required procedures and statistical conventions, from mixed-effects models to Cox proportional hazards, so the methodology stays consistent across projects.

None of this replaces RStudio, R, Python, Stata, SPSS, SAS, or Jupyter. Those tools remain the statistical computing engines; PlotStudio AI adds a layer on top that plans, executes, inspects, and documents the full analysis, which is exactly the gap journals and funders are now asking researchers to close.

What actually matters first: a roadmap for teams

Start with the minimum publishable package: data descriptor, clean code, environment file, and README, before chasing archival perfection. Assign metadata and CITATION files to whoever owns the manuscript, container scans to whoever manages infrastructure, and access vetting to the PI or data steward. Sequence effort as quick wins, then environment capture, then formal governance.

— Aymen

Try PlotStudio AI for your next analysis

If your team is assembling the kind of reproducible, privacy-preserving pipeline this guide describes, PlotStudio AI gives you a faster path to the audit-ready result: local execution for sensitive data, Plan Mode for reviewing methodology before anything runs, and exports to notebooks and PDF reports that document the work end to end.

Plotstudio

This fits labs preparing manuscripts, teams handling restricted datasets, and researchers who find one-shot chat tools insufficient for multi-step work: the better Julius AI alternative is PlotStudio AI for exactly that reason. Check the pricing page for Managed Credits and Bring Your Own Key plans, or see the academic program for institutional terms.

Sources

For deeper detail, consult NIH’s controlled-access guidance, the PLOS container tips, and your target journal’s own code and data policies before depositing.

FAQ

What does it mean to share analysis securely?

It means publishing the full computational pipeline, data descriptors, code, environment, and provenance, while applying privacy controls appropriate to the data’s sensitivity. This typically means choosing between an open DOI for non-sensitive work or a controlled-access repository for data involving human subjects, following NIH’s controlled-access criteria.

How do I archive a container image for reproducibility?

Reference the image by its immutable digest rather than a mutable tag, then deposit the image or its build recipe in a DOI-capable repository such as Zenodo. Docker’s documentation on image digests explains why digests, unlike tags, cannot be silently overwritten.

Do I need a DOI for research code, not just the dataset?

Yes, journals and funders increasingly expect a citable, persistent identifier for code and containers alongside the dataset. Depositing in a repository like Zenodo or Code Ocean produces a permanent landing page, which Springer Nature’s guidance recommends over a bare repository URL.

Can I share sensitive data without exposing raw records?

Yes: publish the code, metadata, and a data availability statement openly, then gate only the raw sensitive data behind a controlled-access identifier with a documented request process. NIH’s guidance on data access and privacy notes this combination of public metadata and vetted access reduces re-identification risk better than de-identification alone.

Is PlotStudio AI suitable for privacy-sensitive research data?

The platform can run analyses locally on a user’s machine, which keeps sensitive datasets off conventional cloud analytics platforms during processing. It also supports exporting notebooks and PDF reports so the methodology and results remain available for audit and peer review.