Ill Effects of College Scattergrams

When looking at a map, Mark Monmonier encourages users to be critical of the subjective influences that have shaped it, emphasizing the undue respect maps receive as compared to other data visualizations. Specifically, Monmonier states that “maps, like numbers, are often arcane images accorded undue respect and credibility,” reflecting his frustration with individuals’ lack of skepticism (Monmonier 3). Well, Monmonier may have been disappointed with my habit of obsessing over college scattergrams during my junior and senior years of high school.

https://www.cappex.com/page/collegeCenter/scattergramStandAlone.jsp;jsessionid=EC6AE4ABF835738BEA2A1BC9B2F08D94.tomcat-main01?id=198385&collegeID=198385&collegeName=Davidson%20College&isForProfitCollege=false

For those who do not know, college scattergrams “are collected data points graphed to show the GPA and test scores of applicants to the college, indicating their accepted pool of students in a visual form” (McNamara). To be honest, I used to scour through these scattergrams for hours, switching from college to college on sites like Cappex , ultimately affecting my decision to apply to certain schools based on the probability that I would be accepted. I fell prey to those who designed these scattergrams, who stripped me of my identity beyond my GPA and test scores. Similarly, scattergrams do not capture the importance of other factors that determine college acceptance, such as extracurriculars, legacy status, college essay, teacher recommendations, etc.. Monmonier would also criticize how rejections are seen as red dots, while acceptances are seen as green, which “mislead the map viewer” into thinking they are an ‘other’ or not worthy enough to be a part of the green dots (Monmonier 3). Therefore, this “visual reduction” of a student negatively impacts their self worth by making it seem as though test scores and GPA are the most prized attributes colleges focus on (Manovich 38). In analyzing data visualizations, we should be mindful of how they only offer a limited and biased view of the bigger picture.

Additional Sources Used:

https://www.cappex.com/page/collegeCenter/scattergramStandAlone.jsp;jsessionid=EC6AE4ABF835738BEA2A1BC9B2F08D94.tomcat-main01?id=198385&collegeID=198385&collegeName=Davidson%20College&isForProfitCollege=false

http://threeforfreecollege.com/2013/09/20/the-truth-about-gpa-test-scores-and-scattergrams/

In Case You Ever Wondered Why Cartographers Don’t Go To Davidson

Mark Monmonier’s article “How to Lie with Maps” draws a lot of parallels with the Lev Manovich article that we read in class last week, highlighting the fact that data in and of itself is nowhere near as important as the human ability to analyze and interpret data. Maps are unable to perfectly represent the region that they are showing, because the cartographer is forced to superimpose a 3-d world onto a 2-d sheet of paper. This goes hand in hand with some of the problems that we have when we obtain data, because it may have been filtered in a certain way that we are not aware of, which could completely change the results of the analysis.

Topographic Map

The example that Monmonier uses with maps is that map users usually trust the benevolence of the mapmakers , which may or not be misplaced. Because cartographers are not licensed, just about anyone can make a map, which means that they might not always be accurate. If someone is aware of this, they might look at the map with a more critical perspective, compare it to the world around them, and then note some of the mistakes that may be present in the map. However, if an individual is unaware of this, they may assume the map is correct, and any problems they experience while using the map must be due to their own personal mistakes. The individual who views the map with a more critical eye will eventually be the one who walks away with a better understanding of how the map he is holding relates to his surroundings. Likewise, in our data analysis, we should be sure to investigate all of the assumptions that the data set uses, as well as other relevant background information in order to make our analysis of the data set be thorough, which will eventually allow us to make the most accurate analysis possible.

Image URL: http://upload.wikimedia.org/wikipedia/commons/2/22/Turkish_Van_Cat.jpg Continue reading In Case You Ever Wondered Why Cartographers Don’t Go To Davidson

Sports Visualization: Is it Useful?

Spurs_Thunder_team_shape_251x140

A common method that coaches tell their players in basketball is to visualize their shot going into the hoop before the action takes place. This is a trick that can help a shooter’s confidence. But are their other ways that sports data can be “visualized” for human use?

This week’s reading by Lev Manovich analyzed the usefulness of data visualization. Manovich provides a rough definition of infovis “as a mapping between discrete data and a visual representation.” The question that arises out of this definition is how accurately does a visual embody the raw data? The answer varies. Manovich mentions data reduction as well as the use of special variables as two main characteristics of infovis. These methods effectively smooth the data in order to represent key aspects to the viewers.

Although useful, I think that these two characteristics of visual analysis can have a tendency to digress from a dataset’s overall meaning. They do offer the viewer an opportunity to make sense of the datapoints, but by omitting aspects, the individualized data could lead to different results. For example, a graph can show a player’s field goal percentage in several games, but the analysis will leave out variables such as whether the shot was contested, minutes played, and the location of the shot.

A new system of NBA tracking has been implemented that tracks the movement of players 25 times per second. This new form of data allows for far more advanced statistics concerning touches, rebound opportunities, drives, and catch and shoots. The system also allows viewers to see video as well as movement animations to form a more complete level of analysis. Manovich would characterize this type of data as “direct visualization.” In this fashion, the NBA tracking system uses actual images and video to make a new depiction of the data. Will this new way of visualization change the way players and coaches think about stats? No way to tell now, but I encourage you to check out the site.

Data is the Oxygen to Machine Learning’s Fire

cat detection
YouTube’s Conception of “Cat”

This past week’s discussion of Big Data seems incomplete without mentioning machine learning. For the uninitiated, machine learning is a method of computing that “trains” algorithms to achieve a desired result using large amounts of data. While machine learning has been around since 1959, it has only recently come back into fashion as businesses are waking up to its potential in the age of Big Data.

MLB post 11 - Image 1.JPG-550x0

The key ingredient in effective machine learning applications, from FaceBook to Baidu, is data. Lots of data. Because of their scale, these companies are able to gather trillions of data points for everything, including individuals’ emotions, shopping habits, facial features and much more. To a human, or even an army of humans, one trillion pictures of peoples’ faces would yield little utility. But, to an elite team of computer scientists, such pictures allow them to construct systems that can recognize identity and emotion more accurately, and at exponentially greater scale, than human beings.

The paradigmatic shift from silos of individual networks to aggregated data and computing resources known as cloud computing brings with it greater efficiency and more data. Joseph Sirosh, Microsoft’s VP of Machine Learning, was recently interviewed by the cloud and data experts at GigaOM. Sirosh explains that Microsoft is rapidly transitioning from an operating system provider to a cloud provider with expertise in big data and machine learning. He goes so far as to state that computing itself is less important than the data that it provides:

“I think you should even first ask, ‘How big is the world of data to computing itself?’” he said. “I would say that in the future, a huge part of the value being generated in the field of computing . . . is going to come from data, as opposed to storage and operating systems and basic infrastructure. It’s the data that is most valuable.”

Microsoft has made its billions by providing software and services. Pivoting the business model of a $360B company is a herculean task; it’s safe to say that Microsoft and every other major player in technology wouldn’t be chasing desperately after data science if it wasn’t a huge deal. Data is often described as the new oil of the 21st century – machine learning is the new refinery.

Continue reading Data is the Oxygen to Machine Learning’s Fire

Procrastination Nation?

Ever since I was a kid, procrastination has been an awful habit of mine. Whether its waiting until the last minute on an assignment or turning it in well after the deadline, I’ve done it all. So for this weeks observation, I decided to look outside the classroom to see if any of my fellow classmates share my love for procrastination based on their posting times for the course blog. Because the “responders” often comment on posts, it was difficult to track their posting times. Thus, I only looked at the posting times for the “readers” the past 4 weeks and the “observers” the past 3 weeks (my group not included). According to our syllabus, the due time for the readers is 10pm on Monday before class and the due time for the observers is 5pm on Friday. The posts fall into 5 categories: 12+ hours before the deadline, 12-7 hours before, 6-4 hours before, 3-0 hours before, and past the deadline. The data is as follows:

ReadersObservers

According to this data, it can be assumed that the Reader groups share my addiction to pushing assignments off, as a majority of the groups posted between 3-0 hours before the deadline. The Observers however, seem to favor posting past the 5pm deadline. This could be due to the fact that the assignment is due on Friday afternoon and the weekend is an inevitable force drawing students away from work. Nonetheless, it was interesting to learn how the posting times differed between the Readers and the Observers over the past few weeks.

People’s Ticks

While observing the class this week, I chose to pay attention to people in the class’s “Ticks” or gestural habits. The two main areas I analyzed were upper body ticks (i.e., fiddling with their hands) and lower body ticks (i.e., heal tapping or knee jiggling). I also observed if people partook in both upper and lower body ticks.

Tuesday- 20 students in class

Lower Body Tick- 11

Upper Body Ticks-5

Both-5

Probability of Lower Body Tick- 55%

Probability of Upper Body Tick- 25%

Probability of Both- 25%

Thursday-21 Students in Class

Lower Body Tick- 7

Upper Body Tick- 5

Both-3

Probability of Lower Body Tick-33.3%

Probability of Upper Body Tick-23.8%

Probability of Both-14.2%

 

These findings were interesting to me because it became clear that many more students tend to have lower body ticks. Also, that people who had upper body ticks were very likely to have a lower body tick in addition to their upper body tick.

My estimation for their being less students with ticks on Thursday is because we did more group work than we did on Tuesday. I believe that group work makes for a more relaxed atmosphere, leading to less reasoning for students to engage in nervous ticks.

 

 

 

 

Entering and Leaving Class

This week I wanted to examine how the class filled up by table, how many jackets were on seat-backs, and how the class exited.

On Tuesday, between the snow and some of the sports teams missing we had 20 students, and the picture below shows how many people sat at each table (the number below the square) and the order that in which the tables filled up (the number in the square).

Tuesday:Screenshot (3)

On Thursday we had a few more people with 22 students. Image follows the same rules as above.

 

Thursday:Screenshot (4)

I was also interested to see how many jackets would be placed on seat-backs for the “snow day” vs. non-snow day.

Tuesday: 8 students had jackets on chair-backs.

8/20= 40% of people

Thursday: 12 students had jackets on chair-backs.

12/22= 54.5% of people

The perception of a snow day could have put people in a mind set that told them they were cold and so more people kept their jacket on.

 

When exiting the room the class has two options: they can leave what I call office side (side facing the front of the library) or writing center side (facing the back of the library).

Tuesday: 16 students left library side.

16/20= 80% of people

Thursday: 13 students left library side.

13/22= 59.1% of people

These numbers could very as people may not have had a class to rush off too on Tuesday, but did on Thursday.

Data about choices: Does it affect how we choose?

During Tuesday’s class, Dr. Sample played a video clip advertising the new Walking Dead video game, which was relevant because of how the game kept track of individual decisions and compiled statistics aggregating all of its users’ choices. Forced to make heat-of-the-moment decisions—“Do I save Ben, or let him fall to his death?”—users were likely to choose options they would later second-guess. The aggregated stats gave context for those split-second decisions, either comforting those in the majority (“I felt bad for letting him die, but at least 79% of users did the same thing”) or compounding regret for the few (“Wow, that really was stupid to try and save him”). Generally, people find comfort in choosing with the majority, although there are certainly rogues out there who would intentionally take the less-beaten path.

tourney pickemWhat interests me is the difference between how people choose when given the stats as opposed to when it’s simply a blind choice—particularly in the context of a NCAA tournament bracket. Every March, thousands of people join online bracket-picking tournaments, often relying on picking the favorites in each first-round matchup (especially for games involving schools no one’s ever heard of). I imagine that, rather than blindly guessing, most people would lean toward the team that has, say, 71% of users picking them, rather than the underdog with 29% on their side. What intrigues me about this is the possibility of a snowball effect, where people disproportionately favor the team with 52% support, which pushes the number higher to 53%, which makes users even more likely to pick them, and so on. And then there are the aforementioned rogues, who intentionally pick an underdog they know nothing about simply because of the thrill of contradicting the mainstream opinion. Once the data on users’ choices is available to the users themselves, decisions can quickly change—in video games, bracket pools, and surely other fields as well.

Image credit: Kawakami, Mark. “Using YUI 3 to Build the Yahoo! Sports Tourney Pick’em Game.” 19 Mar. 2010. Web. Accessed 18 Feb. 2015. <http://yuiblog.com/blog/2010/03/19/tourney-pickem/>

Big Data: positive impact or negative impact?

Big-Data-Blog-Image

The argument becomes more and more contentious between customers actively aggregating data profile and being passively collected personal information. The Target’s case put this issue in a more controversial and extreme situation. Since Target’s effort to speculate women’s pregnancy reached a very deep level of personal privacy that most people wouldn’t feel comfortable being asked about by strangers. However, on the other side, while Target is trying to make profit and solicit customers, it is also bring much convenience and promotions to the targeted women customers. Just as many other e-service-based companies like Netflix and Google are doing, most of advertisements or movie recommendations are the result of Big Data analysis. In most cases, we all benefit from the service brought by Big Data technology.

What can also be a backfire other than privacy violation? As I have talked about in the data critique, the minority population might be neglected through the process of Big Data analysis. Every 5 years, American Census Bureau collects data from various aspects like family, work, and education. With the database, the bureau tries to make estimates which can apply to the entire population. However, from the estimates, we can easily recognize the minority groups from the majority. Thus, it is more likely the private companies would design their products for the majority based on the estimates. Is Big Data hurting the minority groups? What should we do to improve? Perhaps we can predict that in the future, Big Data is able to provide the basis for a perfect market which hurts no minorities.

Picture from: Singh, Tarry “Big Data Is The Future Of Digital Marketing” WordPress. July 26th 2014. Web. Feb. 17th 2015. (http://tarrysingh.com/2014/07/big-data-is-the-future-of-digital-marketing/ )