October 2018
Intermediate to advanced
472 pages
10h 57m
English
Data that we apply to our models are representations of the real world. This is the fundamental truth that unites computational linguistics and CV. With respect to CV, we need to remember that 2D images represent a 3D world, in the same way that video represents 4D, with the added aspects of time and movement. Recalling this obvious fact lets us ask ever more interesting questions and develop deep learning technologies with increasing utility. Our hypothetical use case was to enable visual effects specialists to easily estimate the pose of actors (particularly the shoulders, neck, and head) on frames of a video. Our task was to build the intelligence for this application.
We successfully ...
Read now
Unlock full access