Chapter 86. Understanding the Ways Different Data Domains Solve Problems
Matthew Seal
Technology organizations commonly develop parallel tracks for multiple data concerns that need to operate in tandem. You often get a mix of teams that include data engineering, machine learning, and data infrastructure. However, these groups often have different design approaches, and struggle to understand the decisions and trade-offs made by their adjacent counterparts. For this reason, it’s important for the teams to empathize with and understand the motives and pressures on one another to make for a successful data-driven company.
I have found that a few driving modes of thought determine many initial assumptions across these three groups in particular, and that knowing these motives helps support or refute decisions being made. For example, data science and machine learning teams often introduce complexity to tooling that they develop to solve their problems. Most of the time, getting a more accurate, precise, or specific answer requires adding more data or more complexity to an existing process. For these teams, adding complexity is therefore often a reasonable trade-off to value. Data scientists also tend to be in an exploratory mode for end results, and focusing on optimizing the steps to get there only slows that exploration when most things tried aren’t kept anyway.
For data infrastructure, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access