Spreadsheets
.xlsx and .xls workbooks open as a read-only grid. A workbook with several sheets shows a tab for each. Formula cells show the result the workbook saved, never the formula. The header gives the sheet’s size, for example “1,250 rows, 14 columns”. The grid shows up to 1,000 rows and 100 columns per sheet. A larger sheet reads: “Preview limited to 1,000 rows and 100 columns.” “This spreadsheet could not be opened. It may be encrypted or damaged.” BioBase couldn’t read the workbook. An .xlsx workbook saved with a password reads “This file is password protected. Remove the password in the application that created it, then open the unprotected copy here.”CSV and TSV
.csv and .tsv files open in the editor as a table, with the first row as the column headers and every row of the file. The table splits columns on commas or tabs. Select Source to edit the file as text.NumPy arrays
An .npy file shows its header: shape, dtype, element count, memory order (row-major or column-major) and data size. The viewer doesn’t read the values. An .npz archive lists up to 200 arrays with each one’s shape and dtype. The header reads “NumPy archive, N arrays”. If a member uses a compression BioBase can’t read, the viewer lists its path and size and reads: “Some arrays in this archive use a compression this preview does not read. Their paths and sizes are listed above.”HDF5 and AnnData
.h5, .hdf5 and .h5ad files list their groups and datasets, with each dataset’s shape and dtype. The header reads “HDF5 container, N entries”. The viewer lists the first 400 entries and reads “Listing the first 400 of N entries.” for a larger file. It doesn’t show dataset values or attributes. An AnnData file reads Annotated single-cell matrix, headed by the cell and gene counts, for example2,700 cells × 32,738 genes. It lists the cell annotations, gene annotations, embeddings and layers.
R objects
.rds, .RData and .rda files show the header of the top-level object: its type and length, the R version that wrote it, the earliest R version that reads it, the serialization format and the compression. Files saved uncompressed or with gzip, thesaveRDS default, open. Files saved with xz or bzip2 compression don’t.
Column labels and classes come after the object’s elements in the file, so the viewer doesn’t show them. Load the file in R to see them.
Python pickles
.pkl, .pickle and .joblib files are read without being loaded. The viewer reads: “A pickle is a program, so BioBase reads its instructions without running them. Load it in your own analysis session when you trust its source.” It lists the classes the pickle would construct, for examplepandas.core.frame.DataFrame, up to 40. It also shows the pickle protocol, the instruction count, the dictionaries, lists, tuples and sets it builds, and the file size. A truncated file reads: “The stream ends before its final instruction, so this file is incomplete.”
