Hashtags and Graphs

This post will focus on the post by Peter Saunders about the hashtags. I absolutely agree with the point that hashtags sorted various information into certain categories and “handle the dirty work of aggregating” for the users. It is really interesting that Peter points out the inefficiency caused by the misspelling of hashtags. If you really think about how hashtags work, you will keywords, or tags, used in some blog websites including WordPress work the exactly same way. If you put the network into graphs, you might be able to see tags as endpoints that link to the endpoints on the other side, which are the posts. It resonates with the other reading, “Graphs” from Networks, Crowds, and Markets: Reasoning about a Highly Connected World by David Easley and Jon Kleinberg.

For instance, on instagram, you post a photo with #dataculture. The hashtag serves as a hyperlink between the category and the photo and thus the photo is the endpoint. The other way around, if you search the hashtag #dataculture, the hashtag becomes an endpoint with leads to the category. Just like the graph shown below:

2000px-Goldner-Harary_graph.svg

If you have multiple of photos with multiple hashtags, the graph will go more complicated. And then with a large user group, the graph goes even crazier. This is where Cloud technology came in.

Graph_with_Preferential_attachment

LOD_Cloud_Diagram_as_of_September_2011

To draw a conclusion for this response, the hashtags are certainly inefficient sometimes because of the misspelling, but, instead of creating different databases, the hashtags create multiple hyperlinks to sort information into categories. Comparatively, hashtags are very efficient.

 

Graphs credit to:

http://en.wikipedia.org/wiki/Planar_graph

http://en.wikiversity.org/wiki/Web_Science/Part2:_Emerging_Web_Properties/Emerging_Structure_of_the_Web

http://en.wikipedia.org/wiki/Linked_data

Harness the Power of the #Hashtag

In the third chapter of his book Off the Network: Disrupting the Digital World, Ulises Ali Mejias describes how computers function together to form a network, defined as “a system of linked elements or nodes” (37). Digital networks, Mejias writes, use a combination of human and machine to assume social agency through social tagging systems. The concept may seem abstract, but for anyone who uses social media such as Twitter in their daily lives, tagging works as a beautifully simple way to sort through and classify information on a particular topic.

Perhaps what makes Twitter hashtags so appealing to its users is the fact that it operates as a “folksonomy” with few, if any, rules for what can and can’t be a hashtag. The relatively lawless nature of hashtags comes from the fact that “the negotiation of meaning during the process of classification is delegated from humans to code” (50). Mejias portrays this feature as “perhaps [the] greatest weakness” of social tagging systems, but I argue that any inefficiencies caused by the computer regulating the system are far outweighed by their benefits.

Yes, it is somewhat annoying that the same topic could take on multiple hashtags based on a misspelling or the choice to omit “the,” as in the example below.

twitter capture 3.9

Two people tweeting about the same thing won’t have their tweets sorted into the same “pile” because of their slightly different hashtag choice, but the inefficiency is minor and ultimately insignificant. What makes Twitter hashtags so fascinating is how they can provide a seamless link between celebrities and regular folks. When NASCAR driver Brad Keselowski—who has over 500,000 followers—wants to do an impromptu Q&A session with his fans, all he needs to do is harness the power of the hashtag, and computers handle the dirty work of aggregating all the submissions for him. In my view, it’s well worth a few small sacrifices in order to gain the “individual freedom and scale of access that only the internet can provide” (52).

Social Computing, Social Behavior, Algorithms, and Physics

In Mejias’ chapter on “Computers as Socializing Tools” he sites the critique of social computing, a field focused on the modeling of social behavior, that in some way it simplifies human behavior by “dumbing it down” to a level which a computer can understand it. The creation of artificial intelligence or games like the Sims demonstrate attempts to recreate human life in the digital realm. These attempts so far have fallen short in recreating a 100 percent realistic picture. Therefore, the question remains can computerized algorithms, however complex, ever truly model how we behave? The answer depends who you ask. Critiques such as those mentioned by Mejias might challenge the plausibility. However, some physicists may argue otherwise. Theoretically, if a unified theory of physics such as string theory were ever to be proven, any concept of free will would be nullified. This would mean our brains operate no differently than computers, collecting stimuli and working through algorithms to arrive at an appropriate action. If this is truly the case, then there is no reason computers could not accurately model human behavior. The only obstacle to this goal would be understanding the complex algorithms of the human mind, which is certainly a daunting task.

Expanding the “Small World Phenomenon”

David Easley and Jon Kleinberg’s chapter on graphs provides an instructional background on the mechanics of visual aids. Their explanations on distances between nodes, paths, and connectivity, however, read as overly specific. Developing “basic network properties in a unifying language” (23)  for graphs sounds unnecessary to a common reader because the very point of graphs is to make information accessible without the need for complicated directions. Visuals are so prevalent in our daily life that directions are unnecessary. Even seemingly private information made available to us with visual aids, like our Facebook data, is immediately digestible. Reconsidering how graphs are established, however, allows us to find similarities between visual aids that may be unexpected.

Through their discussion on social media, Easley and Kleinberg hinge upon a link between visual maps and connections with friends online. They introduce a concept named the “small-world phenomenon— the idea that the world looks “small” when you think of how short a path of friends it takes to get from you to almost anyone else” (35). The Guardian’s article on internet privacy opened our eyes to how closely linked humans are- even people who don’t know each other. I would argue that other graphs also work to shrink distant locations, and apply the idea of social network analysis to physical geography. Easley and Kleinberg use airline routes and subway maps to illustrate the graph’s ability to document “direct connections.” Shrinking one-hundred city blocks into twenty orderly dots makes an urban environment look compact. I can travel across an entire city by making a few stops along this lined? A flight from Atlanta to Tokyo, a fifteen hour flight across an entire ocean and continent, can be reduced to a two-page spread in a magazine. That journey looks pretty simple. As previously mentioned, graphs are incredibly accessible and benefit from not needing the mechanical directions provided by Easley and Kleinberg in their chapter. One thing graphs lack, however, is a sense of context. While they may be easy to read, their simplicity may imprint “the small-world phenomenon” on geography as well as people. Accessibility comes with a cost; whether or not such a cost is positive or negative is up for debate.

Vancouver and Amsterdam are only a magazine page away

 

Meaning Behind a Digital Network

Through technology, we choose who we associate ourselves with. The association with someone would create an edge between nodes. The edge forms a network that connects two nodes together where the nodes of family, friends, and associates are now connected to both nodes. For example, Facebook is a digital network where two people can become friends on Facebook. Now, the friends of these two individuals are connected through a path that are formed by two individuals, unless the friends already have mutual friends. Ulises Ali Mejias states in Off the Network how “to ‘Friend’ in a social network that establishes a correspondence between two records of data (47). Every day, Facebook users are being connected together as one can find Facebook to consist of a giant component that encompasses every Facebook user being connected through edges. Even if one creates a new Facebook account, then they will soon friend someone else that will tie them into the giant component. Yet, the downfalls behind the giant component of Facebook include disruptive forces like a person impersonating a real-world friend that could disrupt the connection between you and your friends.

Furthermore, through Facebook, we witness various Facebook users discover their friends to have hidden mutual friends that may have not been discovered through technology. David Easley and Jon Kleinberg states this idea, small-world phenomenon, where the world looks “small” when you think of how short a path of friends it takes to get from you to almost anyone else (35). The downside from this phenomenon is it doesn’t mean you’re socially close to them.

Facebook friends
http://socialmediaiseasy.blogspot.com/2012/11/learn-how-to-use-facebook-basics.html

David Easley and Jon Kleinberg, “Graphs” from Networks, Crowds, and Markets: Reasoning about a Highly Connected World (2010)

Ulises Ali Mejias, “Computers as Socializing Tools” from Off the Network: Disrupting the Digital World (2013)

Data is the Oxygen to Machine Learning’s Fire

cat detection
YouTube’s Conception of “Cat”

This past week’s discussion of Big Data seems incomplete without mentioning machine learning. For the uninitiated, machine learning is a method of computing that “trains” algorithms to achieve a desired result using large amounts of data. While machine learning has been around since 1959, it has only recently come back into fashion as businesses are waking up to its potential in the age of Big Data.

MLB post 11 - Image 1.JPG-550x0

The key ingredient in effective machine learning applications, from FaceBook to Baidu, is data. Lots of data. Because of their scale, these companies are able to gather trillions of data points for everything, including individuals’ emotions, shopping habits, facial features and much more. To a human, or even an army of humans, one trillion pictures of peoples’ faces would yield little utility. But, to an elite team of computer scientists, such pictures allow them to construct systems that can recognize identity and emotion more accurately, and at exponentially greater scale, than human beings.

The paradigmatic shift from silos of individual networks to aggregated data and computing resources known as cloud computing brings with it greater efficiency and more data. Joseph Sirosh, Microsoft’s VP of Machine Learning, was recently interviewed by the cloud and data experts at GigaOM. Sirosh explains that Microsoft is rapidly transitioning from an operating system provider to a cloud provider with expertise in big data and machine learning. He goes so far as to state that computing itself is less important than the data that it provides:

“I think you should even first ask, ‘How big is the world of data to computing itself?’” he said. “I would say that in the future, a huge part of the value being generated in the field of computing . . . is going to come from data, as opposed to storage and operating systems and basic infrastructure. It’s the data that is most valuable.”

Microsoft has made its billions by providing software and services. Pivoting the business model of a $360B company is a herculean task; it’s safe to say that Microsoft and every other major player in technology wouldn’t be chasing desperately after data science if it wasn’t a huge deal. Data is often described as the new oil of the 21st century – machine learning is the new refinery.

Continue reading Data is the Oxygen to Machine Learning’s Fire

Data about choices: Does it affect how we choose?

During Tuesday’s class, Dr. Sample played a video clip advertising the new Walking Dead video game, which was relevant because of how the game kept track of individual decisions and compiled statistics aggregating all of its users’ choices. Forced to make heat-of-the-moment decisions—“Do I save Ben, or let him fall to his death?”—users were likely to choose options they would later second-guess. The aggregated stats gave context for those split-second decisions, either comforting those in the majority (“I felt bad for letting him die, but at least 79% of users did the same thing”) or compounding regret for the few (“Wow, that really was stupid to try and save him”). Generally, people find comfort in choosing with the majority, although there are certainly rogues out there who would intentionally take the less-beaten path.

tourney pickemWhat interests me is the difference between how people choose when given the stats as opposed to when it’s simply a blind choice—particularly in the context of a NCAA tournament bracket. Every March, thousands of people join online bracket-picking tournaments, often relying on picking the favorites in each first-round matchup (especially for games involving schools no one’s ever heard of). I imagine that, rather than blindly guessing, most people would lean toward the team that has, say, 71% of users picking them, rather than the underdog with 29% on their side. What intrigues me about this is the possibility of a snowball effect, where people disproportionately favor the team with 52% support, which pushes the number higher to 53%, which makes users even more likely to pick them, and so on. And then there are the aforementioned rogues, who intentionally pick an underdog they know nothing about simply because of the thrill of contradicting the mainstream opinion. Once the data on users’ choices is available to the users themselves, decisions can quickly change—in video games, bracket pools, and surely other fields as well.

Image credit: Kawakami, Mark. “Using YUI 3 to Build the Yahoo! Sports Tourney Pick’em Game.” 19 Mar. 2010. Web. Accessed 18 Feb. 2015. <http://yuiblog.com/blog/2010/03/19/tourney-pickem/>

Big Data: positive impact or negative impact?

Big-Data-Blog-Image

The argument becomes more and more contentious between customers actively aggregating data profile and being passively collected personal information. The Target’s case put this issue in a more controversial and extreme situation. Since Target’s effort to speculate women’s pregnancy reached a very deep level of personal privacy that most people wouldn’t feel comfortable being asked about by strangers. However, on the other side, while Target is trying to make profit and solicit customers, it is also bring much convenience and promotions to the targeted women customers. Just as many other e-service-based companies like Netflix and Google are doing, most of advertisements or movie recommendations are the result of Big Data analysis. In most cases, we all benefit from the service brought by Big Data technology.

What can also be a backfire other than privacy violation? As I have talked about in the data critique, the minority population might be neglected through the process of Big Data analysis. Every 5 years, American Census Bureau collects data from various aspects like family, work, and education. With the database, the bureau tries to make estimates which can apply to the entire population. However, from the estimates, we can easily recognize the minority groups from the majority. Thus, it is more likely the private companies would design their products for the majority based on the estimates. Is Big Data hurting the minority groups? What should we do to improve? Perhaps we can predict that in the future, Big Data is able to provide the basis for a perfect market which hurts no minorities.

Picture from: Singh, Tarry “Big Data Is The Future Of Digital Marketing” WordPress. July 26th 2014. Web. Feb. 17th 2015. (http://tarrysingh.com/2014/07/big-data-is-the-future-of-digital-marketing/ )

Invisible Observations: Week of February 2nd

Tasked with making observations and gathering data for two DIG 210 classes, February 3rd and 5th, I sought to try something unconventional. Rather than focus on visible characteristic(s), I thought I might gather data about the auditory aspect of our class. Given that we exchange knowledge in our class primary over the medium of sound, through our voices, I thought it might be interesting to transpose the sound waves that reverberated through Studio D of the Davidson library from 1:40pm to 2:55pm on those two days into a static 2D image that we can view and make observations about.

Sound Waves Captured on February 3rd
Sound Waves Captured on February 3rd
Sound Waves Captured on February 5th
Sound Waves Captured on February 5th

There are several limitations to this method of observation. One is the fact that my computer’s microphone was the sole conduit for generating data; being in one place and not designed with high-fidelity audio capture in mind, the data that my computer generates are limited in accuracy. Sounds that I or others at my table made register more prominently than equally loud sounds made by others. These data paint a picture from the vantage point of my computer, which on both occasions I tried to position as close to the center of the room as possible. Finally, my skills in audio analysis are extremely limited; with more expertise and better software, I would be able to provide many more statistics that might allow us to glean more information from these data.

The blue lines vertically indicate amplitude of the sound waves over the horizontal axis of time. We notice by comparing the two that the February 5th class was, on average, louder than the February 3rd class – given that a guest speaker spoke to the class through many speakers on the 5th, whereas Dr. Sample primarily lectured using only his vocal cords on the 3rd, this makes sense. We notice in both graphs more peaks toward the end of the class as opposed to the beginning, which could be the result of many factors – perhaps the conversation becomes more heated and interesting when the class is more involved in the material?

Note: I have chosen to not make the actual sound files available for a variety of reasons.

Week 4 Data Collection

I wasn’t able to be in class this week, so instead of gathering data in class, I took a look at how our class used the blog. Specifically, I took a look at the frequency of certain words, similar to the State of the Union Adress database we examined.

Data From 2/3: Reader Blog Posts: How many times did Readers use the following words?

Graphics: 2

Works of Art: 2

Photograph: 1

Hemings Count: 13

Jefferson Count: 27

Slavery: 13

Data: 9

Search: 8

The: 118

 Data From 2/5: Kissinger Questions: How many times were the following words used by groups?

Data/database: 2

How: 7

What: 10

Why: 4

Mean: 4

Textplot: 6

Color: 3

Network: 2

Word: 4