Data is the Oxygen to Machine Learning’s Fire

cat detection
YouTube’s Conception of “Cat”

This past week’s discussion of Big Data seems incomplete without mentioning machine learning. For the uninitiated, machine learning is a method of computing that “trains” algorithms to achieve a desired result using large amounts of data. While machine learning has been around since 1959, it has only recently come back into fashion as businesses are waking up to its potential in the age of Big Data.

MLB post 11 - Image 1.JPG-550x0

The key ingredient in effective machine learning applications, from FaceBook to Baidu, is data. Lots of data. Because of their scale, these companies are able to gather trillions of data points for everything, including individuals’ emotions, shopping habits, facial features and much more. To a human, or even an army of humans, one trillion pictures of peoples’ faces would yield little utility. But, to an elite team of computer scientists, such pictures allow them to construct systems that can recognize identity and emotion more accurately, and at exponentially greater scale, than human beings.

The paradigmatic shift from silos of individual networks to aggregated data and computing resources known as cloud computing brings with it greater efficiency and more data. Joseph Sirosh, Microsoft’s VP of Machine Learning, was recently interviewed by the cloud and data experts at GigaOM. Sirosh explains that Microsoft is rapidly transitioning from an operating system provider to a cloud provider with expertise in big data and machine learning. He goes so far as to state that computing itself is less important than the data that it provides:

“I think you should even first ask, ‘How big is the world of data to computing itself?’” he said. “I would say that in the future, a huge part of the value being generated in the field of computing . . . is going to come from data, as opposed to storage and operating systems and basic infrastructure. It’s the data that is most valuable.”

Microsoft has made its billions by providing software and services. Pivoting the business model of a $360B company is a herculean task; it’s safe to say that Microsoft and every other major player in technology wouldn’t be chasing desperately after data science if it wasn’t a huge deal. Data is often described as the new oil of the 21st century – machine learning is the new refinery.

My colleague Scott Patrick’s Procrastination Nation? post would seem apropos – apologies for the delayed response!

Leave a Reply