Data Storage and TrackingπŸ”—

Best PracticesπŸ”—

MSI is a shared use space, so it is important to know the approximate size impact your data (inputs and outputs) will have on the shared environment. When you plan to process a dataset, we recommend running jobs for a typical subject on Tier 1 and finding out how large the outputs are. Multiply that size by the total number of subjects in your dataset and add the size of your inputs for the full impact. You can run this command to check how large your input/output directories are: du -sh --total /input/or/output/path

CDNI data are often used by multiple researchers and analysts. Please document data locations to prevent 'double dipping' on storage space, especially when your work requires more than 1TB of space. For Tier 1 use, submit a storage request form about what kinds of data you will be putting onto Tier 1 storage and why Tier 1 is needed, specifically.

Data storage options:πŸ”—

  • Tier 1 space is limited to 150GB - 20TB per group (depending on the group's allocation). To see what your group allocation is, copy the link https://www-archive.msi.umn.edu/group/<group>/storage into your browser bar with the group name in the path. You can also use the command groupquota -g <group>.

  • S3 (aka Tier 2) storage should be used whenever possible. New MSI users are limited to 5GB of Tier 2 storage. PIs are allocated 120TB of Tier 2 storage. If you need to store new data on the s3, please ask your PI to create a bucket for you. To see how much space is used in a particular bucket, use the command s3cmd du -H s3://<bucketname>/. Access to a S3 storage bucket can usually be granted by any user who already has access to that bucket. You can learn more about the S3 here

Attention

Outputs using ABCD/ABCC data CANNOT BE STORED ON TIER 1. These spaces are not under control of the UMN ABCD designated user credentials (DUC), which is required to access ABCD data. All ABCD outputs should either be stored on the `cdni-nih-bdc` share, or processed in tmp space and synced to the s3.

Data TrackingπŸ”—

  • For CDNI data being stored on Tier 1, documentation of the dataset description, dataset location, dataset size, the primary owner of the data, and the estimated start and end date for use of the necessary storage space should be entered in the Tier 1 Tab of our Data Location Google Sheet. For further information on allowable data on each share, visit the β€œShares” tab on the same spreadsheet.

  • For CDNI data being stored on s3, documentation of the bucket name, dataset description, dataset location, bucket owner, and users who has access to the bucket should be entered in the Tier 2 of our Data Location Google Sheet. It would be helpful to include what pipeline the dataset was processed with and what version in the dataset description.

  • For information on backup and versioning procedures, please refer to the Backup and Versioning PowerPoint Presentation.

Transfering Data from MSI to Local StorageπŸ”—

Your UMN computer has a local storage space that you may want to use for working on specific files. You can transfer files between local and MSI Tier 1 systems via winSCP or FileZilla.

For questions, suggestions, or to note any errors, post a Github issue.