book

Agile Data Science

Name: Agile Data Science
Author: Russell Jurney
ISBN: 9781449326265

by Russell Jurney

October 2013

Beginner to intermediate

175 pages

3h 53m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
Who This Book Is ForHow This Book Is OrganizedConventions Used in This BookUsing Code ExamplesSafari® Books OnlineHow to Contact Us
I. Setup
1. Theory
Agile Big DataBig Words DefinedAgile Big Data TeamsRecognizing the Opportunity and ProblemAdapting to ChangeHarnessing the power of generalistsLeveraging agile platformsSharing intermediate resultsAgile Big Data ProcessCode Review and Pair ProgrammingAgile Environments: Engineering ProductivityCollaboration SpacePrivate SpacePersonal SpaceRealizing Ideas with Large-Format Printing
2. Data
EmailWorking with Raw DataRaw EmailStructured Versus Semistructured DataSQLNoSQLSerializationExtracting and Exposing Features in Evolving SchemasData PipelinesData PerspectivesNetworksTime SeriesNatural LanguageProbabilityConclusion
3. Agile Tools
Scalability = SimplicityAgile Big Data ProcessingSetting Up a Virtual Environment for PythonSerializing Events with AvroAvro for PythonInstallationTestingCollecting DataData Processing with PigInstalling PigPublishing Data with MongoDBInstalling MongoDBInstalling MongoDB’s Java DriverInstalling mongo-hadoopPushing Data to MongoDB from PigSearching Data with ElasticSearchInstallationElasticSearch and Pig with WonderdogInstalling WonderdogWonderdog and PigSearching our dataPython and ElasticSearch with pyelasticsearchReflecting on our WorkflowLightweight Web ApplicationsPython and FlaskFlask Echo ch03/python/flask_echo.pyPython and Mongo with pymongoDisplaying sent_counts in FlaskPresenting Our DataInstalling BootstrapBooting BoostrapVisualizing Data with D3.js and nvd3.jsConclusion
4. To the Cloud!
IntroductionGitHubdotCloudEcho on dotCloudPython WorkersAmazon Web ServicesSimple Storage ServiceElastic MapReduceMongoDB as a ServicePushing data from Pig to MongoDB at dotCloudInstrumentationGoogle AnalyticsMortar Data
II. Climbing the Pyramid
5. Collecting and Displaying Records
Putting It All TogetherCollect and Serialize Our InboxProcess and Publish Our EmailsPresenting Emails in a BrowserServing Emails with Flask and pymongoRendering HTML5 with Jinja2Agile CheckpointListing EmailsListing Emails with MongoDBAnatomy of a PresentationReinventing the wheel?Prototyping back from HTMLSearching Our EmailIndexing Our Email with Pig, ElasticSearch, and WonderdogSearching Our Email on the WebConclusion
6. Visualizing Data with Charts
Good ChartsExtracting Entities: Email AddressesExtracting EmailsVisualizing TimeConclusion
7. Exploring Data with Reports
Building Reports with Multiple ChartsLinking RecordsExtracting Keywords from Emails with TF-IDFConclusion

8. Making Predictions
Predicting Response Rates to EmailsPersonalizationConclusion
9. Driving Actions
Properties of Successful EmailsBetter Predictions with Naive BayesP(Reply | From & To)P(Reply | Token)Making Predictions in Real TimeLogging EventsConclusion
Index
About the Author
Colophon
Copyright

Overview

Mining big data requires a deep investment in people and time. How can you be sure you’re building the right models? With this hands-on book, you’ll learn a flexible toolset and methodology for building effective analytics applications with Hadoop.

Using lightweight tools such as Python, Apache Pig, and the D3.js library, your team will create an agile environment for exploring data, starting with an example application to mine your own email inboxes. You’ll learn an iterative approach that enables you to quickly change the kind of analysis you’re doing, depending on what the data is telling you. All example code in this book is available as working Heroku apps.

Create analytics applications by using the agile big data development methodology
Build value from your data in a series of agile sprints, using the data-value stack
Gain insight by using several data structures to extract multiple features from a single dataset
Visualize data with charts, and expose different aspects through interactive reports
Use historical data to predict the future, and translate predictions into action
Get feedback from users after each sprint to keep your project on track

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781449326890Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills