Chapter 95. Why Data Science Teams Need Generalists, Not Specialists
Eric Colson
In The Wealth of Nations, Adam Smith demonstrates how the division of labor is the chief source of productivity gains, using the vivid example of a pin factory assembly line: “One person draws out the wire, another straightens it, a third cuts it, a fourth points it, a fifth grinds it.” With specialization oriented around function, each worker becomes highly skilled in a narrow task, leading to process efficiencies.
The allure of such efficiencies has led us to organize even our data science teams by specialty functions such as data engineers, machine learning engineers, research scientists, causal inference scientists, and so on. Specialists’ work is coordinated by a product manager, with handoffs between the functions in a manner resembling the pin factory: “one person sources the data, another models it, a third implements it, a fourth measures it,” and on and on.
The challenge with this approach is that data science products and services can rarely be designed up front. They need to be learned and developed via iteration. Yet, when development is distributed among multiple specialists, several forces can hamper iteration cycles. Coordination costs—the time spent communicating, discussing, and justifying each change—scale proportionally with the number of people involved.
Even with just a few specialists, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access