book

Python and HDF5

Name: Python and HDF5
Author: Andrew Collette
ISBN: 9781449367831

by Andrew Collette

November 2013

Intermediate to advanced

148 pages

3h 21m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
Conventions Used in This BookUsing Code ExamplesSafari® Books OnlineHow to Contact UsAcknowledgments
1. Introduction
Python and HDF5Organizing Data and MetadataCoping with Large Data VolumesWhat Exactly Is HDF5?HDF5: The FileHDF5: The LibraryHDF5: The Ecosystem
2. Getting Started
HDF5 BasicsSetting UpPython 2 or Python 3?Code ExamplesNumPyHDF5 and h5pyIPythonTiming and OptimizationThe HDF5 ToolsHDFViewViTablesCommand Line ToolsYour First HDF5 FileUse as a Context ManagerFile Driverscore driverfamily drivermpio driverThe User Block
3. Working with Datasets
Dataset BasicsType and ShapeReading and WritingCreating Empty DatasetsSaving Space with Explicit Storage TypesAutomatic Type Conversion and Direct ReadsReading with astypeReshaping an Existing ArrayFill ValuesReading and Writing DataUsing Slicing EffectivelyStart-Stop-Step IndexingMultidimensional and Scalar SlicingBoolean IndexingCoordinate ListsAutomatic BroadcastingReading Directly into an Existing ArrayA Note on Data TypesResizing DatasetsCreating Resizable DatasetsData Shuffling with resizeWhen and How to Use resize
4. How Chunking and Compression Can Help You
Contiguous StorageChunked StorageSetting the Chunk ShapeAuto-ChunkingManually Picking a ShapePerformance Example: Resizable DatasetsFilters and CompressionThe Filter PipelineCompression FiltersGZIP/DEFLATE CompressionSZIP CompressionLZF CompressionPerformanceOther FiltersSHUFFLE FilterFLETCHER32 FilterThird-Party Filters
5. Groups, Links, and Iteration: The “H” in HDF5
The Root Group and SubgroupsGroup BasicsDictionary-Style AccessSpecial PropertiesWorking with LinksHard LinksFree Space and RepackingSoft LinksExternal LinksA Note on Object NamesUsing get to Determine Object TypesUsing require to Simplify Your ApplicationIteration and ContainershipHow Groups Are Actually StoredDictionary-Style IterationContainership TestingMultilevel Iteration with the Visitor PatternVisit by NameMultiple Links and visitVisiting ItemsCanceling Iteration: A Simple Search MechanismCopying ObjectsSingle-File CopyingObject Comparison and Hashing
6. Storing Metadata with Attributes
Attribute BasicsType GuessingStrings and File CompatibilityPython ObjectsExplicit TypingReal-World Example: Accelerator Particle DatabaseApplication Format on Top of HDF5Analyzing the Data
7. More About Types
The HDF5 Type SystemIntegers and FloatsFixed-Length StringsVariable-Length StringsThe vlen String Data TypeWorking with vlen String DatasetsByte Versus Unicode StringsUsing Unicode StringsDon’t Store Binary Data in Strings!Future-Proofing Your Python 2 ApplicationCompound TypesComplex NumbersEnumerated TypesBooleansThe array TypeOpaque TypesDates and Times
8. Organizing Data with References, Types, and Dimension Scales
Object ReferencesCreating and Resolving ReferencesReferences as “Unbreakable” LinksReferences as DataRegion ReferencesCreating Region References and ReadingFancy IndexingFinding Datasets with Region ReferencesNamed TypesThe Datatype ObjectLinking to Named TypesManaging Named TypesDimension ScalesCreating Dimension ScalesAttaching Scales to a Dataset
9. Concurrency: Parallel HDF5, Threading, and Multiprocessing
Python Parallel BasicsThreadingMultiprocessingMPI and Parallel HDF5A Very Quick Introduction to MPIMPI-Based HDF5 ProgramCollective Versus Independent OperationsAtomicity Gotchas

10. Next Steps
Asking for HelpContributing
Index
About the Author
Colophon
Copyright

Content preview from Python and HDF5

Chapter 2. Getting Started

HDF5 Basics

Before we jump into Python code examples, it’s useful to take a few minutes to address how HDF5 itself is organized. Figure 2-1 shows a cartoon of the various logical layers involved when using HDF5. Layers shaded in blue are internal to the library itself; layers in green represent software that uses HDF5.

Most client code, including the Python packages h5py and PyTables, uses the native C API (HDF5 is itself written in C). As we saw in the introduction, the HDF5 data model consists of three main public abstractions: datasets (see Chapter 3), groups (see Chapter 5), and attributes (see Chapter 6)in addition to a system to represent types. The C API (and Python code on top of it) is designed to manipulate these objects.

HDF5 uses a variety of internal data structures to represent groups, datasets, and attributes. For example, groups have their entries indexed using structures called “B-trees,” which make retrieving and creating group members very fast, even when hundreds of thousands of objects are stored in a group (see How Groups Are Actually Stored). You’ll generally only care about these data structures when it comes to performance considerations. For example, when using chunked storage (see Chapter 4), it’s important to understand how data is actually organized on disk.

The next two layers have to do with how your data makes its way onto disk. HDF5 objects all live in a 1D logical address space, like in a regular file. However, there’s an ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491944981Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Python and HDF5

by Andrew Collette

Chapter 2. Getting Started

HDF5 Basics

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.