book

Python and HDF5

Name: Python and HDF5
Author: Andrew Collette
ISBN: 9781449367831

by Andrew Collette

November 2013

Intermediate to advanced

148 pages

3h 21m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
Conventions Used in This BookUsing Code ExamplesSafari® Books OnlineHow to Contact UsAcknowledgments
1. Introduction
Python and HDF5Organizing Data and MetadataCoping with Large Data VolumesWhat Exactly Is HDF5?HDF5: The FileHDF5: The LibraryHDF5: The Ecosystem
2. Getting Started
HDF5 BasicsSetting UpPython 2 or Python 3?Code ExamplesNumPyHDF5 and h5pyIPythonTiming and OptimizationThe HDF5 ToolsHDFViewViTablesCommand Line ToolsYour First HDF5 FileUse as a Context ManagerFile Driverscore driverfamily drivermpio driverThe User Block
3. Working with Datasets
Dataset BasicsType and ShapeReading and WritingCreating Empty DatasetsSaving Space with Explicit Storage TypesAutomatic Type Conversion and Direct ReadsReading with astypeReshaping an Existing ArrayFill ValuesReading and Writing DataUsing Slicing EffectivelyStart-Stop-Step IndexingMultidimensional and Scalar SlicingBoolean IndexingCoordinate ListsAutomatic BroadcastingReading Directly into an Existing ArrayA Note on Data TypesResizing DatasetsCreating Resizable DatasetsData Shuffling with resizeWhen and How to Use resize
4. How Chunking and Compression Can Help You
Contiguous StorageChunked StorageSetting the Chunk ShapeAuto-ChunkingManually Picking a ShapePerformance Example: Resizable DatasetsFilters and CompressionThe Filter PipelineCompression FiltersGZIP/DEFLATE CompressionSZIP CompressionLZF CompressionPerformanceOther FiltersSHUFFLE FilterFLETCHER32 FilterThird-Party Filters
5. Groups, Links, and Iteration: The “H” in HDF5
The Root Group and SubgroupsGroup BasicsDictionary-Style AccessSpecial PropertiesWorking with LinksHard LinksFree Space and RepackingSoft LinksExternal LinksA Note on Object NamesUsing get to Determine Object TypesUsing require to Simplify Your ApplicationIteration and ContainershipHow Groups Are Actually StoredDictionary-Style IterationContainership TestingMultilevel Iteration with the Visitor PatternVisit by NameMultiple Links and visitVisiting ItemsCanceling Iteration: A Simple Search MechanismCopying ObjectsSingle-File CopyingObject Comparison and Hashing
6. Storing Metadata with Attributes
Attribute BasicsType GuessingStrings and File CompatibilityPython ObjectsExplicit TypingReal-World Example: Accelerator Particle DatabaseApplication Format on Top of HDF5Analyzing the Data
7. More About Types
The HDF5 Type SystemIntegers and FloatsFixed-Length StringsVariable-Length StringsThe vlen String Data TypeWorking with vlen String DatasetsByte Versus Unicode StringsUsing Unicode StringsDon’t Store Binary Data in Strings!Future-Proofing Your Python 2 ApplicationCompound TypesComplex NumbersEnumerated TypesBooleansThe array TypeOpaque TypesDates and Times
8. Organizing Data with References, Types, and Dimension Scales
Object ReferencesCreating and Resolving ReferencesReferences as “Unbreakable” LinksReferences as DataRegion ReferencesCreating Region References and ReadingFancy IndexingFinding Datasets with Region ReferencesNamed TypesThe Datatype ObjectLinking to Named TypesManaging Named TypesDimension ScalesCreating Dimension ScalesAttaching Scales to a Dataset
9. Concurrency: Parallel HDF5, Threading, and Multiprocessing
Python Parallel BasicsThreadingMultiprocessingMPI and Parallel HDF5A Very Quick Introduction to MPIMPI-Based HDF5 ProgramCollective Versus Independent OperationsAtomicity Gotchas

10. Next Steps
Asking for HelpContributing
Index
About the Author
Colophon
Copyright

Content preview from Python and HDF5

Chapter 3. Working with Datasets

Datasets are the central feature of HDF5. You can think of them as NumPy arrays that live on disk. Every dataset in HDF5 has a name, a type, and a shape, and supports random access. Unlike the built-in np.save and friends, there’s no need to read and write the entire array as a block; you can use the standard NumPy syntax for slicing to read and write just the parts you want.

Dataset Basics

First, let’s create a file so we have somewhere to store our datasets:

>>> f = h5py.File("testfile.hdf5")

Every dataset in an HDF5 file has a name. Let’s see what happens if we just assign a new NumPy array to a name in the file:

>>> arr = np.ones((5,2))
>>> f["my dataset"] = arr
>>> dset = f["my dataset"]
>>> dset
<HDF5 dataset "my dataset": shape (5, 2), type "<f8">

We put in a NumPy array but got back something else: an instance of the class h5py.Dataset. This is a “proxy” object that lets you read and write to the underlying HDF5 dataset on disk.

Type and Shape

Let’s explore the Dataset object. If you’re using IPython, type dset. and hit Tab to see the object’s attributes; otherwise, do dir(dset). There are a lot, but a few stand out:

>>> dset.dtype
dtype('float64')

Each dataset has a fixed type that is defined when it’s created and can never be changed. HDF5 has a vast, expressive type mechanism that can easily handle the built-in NumPy types, with few exceptions. For this reason, h5py always expresses the type of a dataset using standard NumPy dtype objects.

There’s ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491944981Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Python and HDF5

by Andrew Collette

Chapter 3. Working with Datasets

Dataset Basics

Type and Shape

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

More than 5,000 organizations count on O’Reilly

Julian F.

Addison B.

Amir M.

Mark W.

You might also like

Python Distilled

Object-Oriented Python

Python Programming Language

Using Asyncio in Python

Publisher Resources

Chapter 3. Working with Datasets

Dataset Basics

Type and Shape

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,and much more.

More than 5,000 organizations count on O’Reilly

Julian F.

Addison B.

Amir M.

Mark W.

You might also like

Python Distilled

Object-Oriented Python

Python Programming Language

Using Asyncio in Python

Publisher Resources

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.