Notice: Function _load_textdomain_just_in_time was called incorrectly. Translation loading for the boldgrid-inspirations domain was triggered too early. This is usually an indicator for some code in the plugin or theme running too early. Translations should be loaded at the init action or later. Please see Debugging in WordPress for more information. (This message was added in version 6.7.0.) in /home/nearth9/public_html/wp-includes/functions.php on line 6170

Notice: Function _load_textdomain_just_in_time was called incorrectly. Translation loading for the heartbeat-control domain was triggered too early. This is usually an indicator for some code in the plugin or theme running too early. Translations should be loaded at the init action or later. Please see Debugging in WordPress for more information. (This message was added in version 6.7.0.) in /home/nearth9/public_html/wp-includes/functions.php on line 6170

Notice: Function _load_textdomain_just_in_time was called incorrectly. Translation loading for the nginx-helper domain was triggered too early. This is usually an indicator for some code in the plugin or theme running too early. Translations should be loaded at the init action or later. Please see Debugging in WordPress for more information. (This message was added in version 6.7.0.) in /home/nearth9/public_html/wp-includes/functions.php on line 6170

Notice: Function _load_textdomain_just_in_time was called incorrectly. Translation loading for the ninja-forms domain was triggered too early. This is usually an indicator for some code in the plugin or theme running too early. Translations should be loaded at the init action or later. Please see Debugging in WordPress for more information. (This message was added in version 6.7.0.) in /home/nearth9/public_html/wp-includes/functions.php on line 6170
Big Data – Earthtech https://1earthtech.com Wed, 02 Sep 2020 02:51:02 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.5 146148357 Artificial Intelligence and Augmented Intelligence: A deeper look https://1earthtech.com/artificial-augmented-intelligence/ https://1earthtech.com/artificial-augmented-intelligence/#respond Fri, 28 Aug 2020 16:26:54 +0000 https://1earthtech.com/?p=1060  

 

Whenever there is a talk of infusing technology, automation and intelligence in to tasks or processes, the solutions proposed fall in to two broad buckets. The first is the conventional use of the term Artificial intelligence, machine learning, deep learning and cognitive learning, which typically work outside human intervention and are largely driven by some kind of an algorithm to achieve its goal. The handling of regular transactions as well as any exceptions are defined within the AI program and the machine doing the job in case of touch and feel jobs like manufacturing or the algorithm producing the result in case of purely online work. The second category is where Artificial Intelligence is applied to work in conjunction with the humans and human intervention and interaction is a critical part of the successful completion of the job. This second type is known as Augmented intelligence or Augmented reality.

 

Another way of looking at what happens with augmented and artificial intelligence is the degree to which algorithms and the programs running the physical machines are expected to make decisions on their own. Is the decision making assisted by humans, controlled by humans or completely out of control of humans?

 

It’s critical to note that while the technologies and fundamentals powering both Artificial intelligence and Augmented intelligence are largely the same, the applications, goals and objectives are objectively different. Simply put, AI creates a human less environment while Augmented Intelligence or Intelligence Augmented as it’s often called, seeks to create an environment for betterment of humans and human endeavors.

 

There’s virtually no major industry where modern AI — more specifically, “narrow AI,” which performs objective functions using data-trained models and often falls into the categories of deep learning or machine learning — hasn’t already affected. That’s especially true in the past few years, as data collection and analysis has ramped up considerably thanks to robust IoT connectivity, the proliferation of connected devices and ever-speedier computer processing.

  • In Manufacturing: AI powered robots work alongside humans to perform a limited range of tasks like assembly and stacking, and predictive analysis sensors keep equipment running smoothly.
  • Healthcare: In the comparatively AI-nascent field of healthcare, diseases are more quickly and accurately diagnosed, drug discovery is sped up and streamlined, virtual nursing assistants monitor patients and big data analysis helps to create a more personalized patient experience.
  • Education: Textbooks are digitized with the help of AI, early-stage virtual tutors assist human instructors and facial analysis gauges the emotions of students to help determine who’s struggling or bored and better tailor the experience to their individual needs.
  • Media: Journalism is harnessing AI, too, and will continue to benefit from it. Bloomberg uses Cyborg technology to help make quick sense of complex financial reports. The Associated Press employs the natural language abilities of Automated Insights to produce 3,700 earning reports stories per year — nearly four times more than in the recent past.
  • Customer Service: Last but hardly least, Google is working on an AI assistant that can place human-like calls to make appointments at, say, your neighborhood hair salon. In addition to words, the system understands context and nuance.

 

Source: McKinsey & Company

 

 

According to a recent publication, of the 9,100 patents received by IBM inventors in 2018, 1,600 (or nearly 18 percent) were AI-related. Here’s another: Tesla founder and tech titan Elon Musk recently donated $10 million to fund ongoing research at the non-profit research company OpenAI — a mere drop in the proverbial bucket if his $1 billion co-pledge in 2015 is any indication. And in 2017, Russian president Vladimir Putin told school children that “Whoever becomes the leader in this sphere [AI] will become the ruler of the world.” He then tossed his head back and laughed maniacally.

 

The single biggest strength of Artificial intelligence is the simple fact that it can do repetitive tasks better than anything else. And the more quantitative, the more objective the job is—separating things into bins, washing dishes, picking fruits and answering customer service calls—those are very much scripted tasks that are repetitive and routine in nature. In the matter of five, 10 or 15 years, they will be displaced by AI.

 

Yet even the most pragmatic AI scientists stress that today’s AI is useless in two significant ways: it has no creativity and no capacity for compassion or love. Rather, it’s “a tool to amplify human creativity.” In other words, sure you can teach AI to make strokes with a brush in a canvas, but an AI driven machine can never in a million years paint a Monalisa.

 

The most advanced robot costing many millions in research dollars and years of work, cannot pickup a tea cup like a 2 year old can. Or millions of sensors on a robot cannot feel taste and smell like humans do. Sure a robot Make millions of calculations in a second, which is many hundreds of times more than the most intelligent human can, sure a robot can discern each smell, or can break down each material in to its atoms and sub-atoms, but it cannot experience the same thoughts, emotions and feelings that human brain can.

 

The second biggest fear around AI stems from the AI and machine learning applications inheriting some of the human biases based on pre-disposed notions and behaviors that humans may code in to or indirectly influence the program.

 

Augmented Intelligence on the other hand doesn’t replace humans with technology rather uses Artificial intelligence to aid humans in doing a job.

 

In recent years, in AI technology rankings in terms of the value they create for businesses, Augmented Intelligence was ranked in second place, just below virtual agents. However, Gartner predicts that “Decision support and AI augmentation will surpass all other types of AI initiatives” creeping into first place this year and then exploding as we reach 2025 becoming around twice as valuable as virtual agents.

 

Just like regular project management work, any organization looking to employ AI needs to first define its use cases and requirements clearly and describe the Big Y, or the big problem it’s trying to solve. Then, they need to define the various aspects of business goals, or the product that they want to build. The next step is to outline data requirements and functionality to solve those business goals. And finally, what’s the ROI or returns the organizations expects. What makes AI and machine or cognitive learning projects unique is that they also need to consider the human-machine dynamic. For example, which part of the chain or which components of the product or solution do they want to hand over entirely to machines to execute, and which parts do they want to retain for their Human Resources? Where machine or a computer algorithm based program is making the decision, is it completely autonomous or is there a human there to monitor? Is The machine or computer program Only responsible to feed data, information or half finished product to human to make the final decisions?  In the scope of these decisions, then, it makes a lot more sense to create an AI role matrix which can evolve over time. This matrix Lays down specifically the kind of role AI or cognitive learning system will play for that particular function or process.

Source: USM Business Systems

 

Getting to true autonomous intelligence or fully independent, cognitive learning powered AI model is proving to be real difficult. Even one single instance of non-compliance or faulty decision making can throw the entire program off tracks.

 

During Tesla’s 2019 Autonomy Day, Tesla CEO Elon Musk said the company is expected to have one million vehicles on the road by the end of 2020 that could function as robotaxis. Though the semi-autonomous Autopilot and Full Self-Driving, or FSD, features are loved by some, others say Musk’s driverless dream is far from becoming a reality. Despite many hundreds of millions in investment and almost a decade of efforts, Tesla’s fully autonomous driving mode is not fully autonomous anymore rather in the best-case scenario is now only viewed as a partial augmented driving mode. In other words, human driver has to control the inputs. And while some of the Tesla’s self-driving features are loved by some, many still view it as unfit for our roads. While the hype around the autonomous driving mode was skyrocketing, and just as Tesla started charging a hefty premium for their self-driving feature, a handful of incidents and accidents changed all that. Scenarios of false-positives, or false-negatives, which are built in to the algorithms powering the programs behind the autonomous driving software, lead to further complications of their own, creating as many problems as they solve. In the best-case scenario, public confidence in self-driving technology is still many years if not decades away.

 

Source: ABC News

 

Watching the Youtube training sessions posted by Waymo, the Alphabet subsidiary in charge of self-driving cars, reveals the concerns with self-driving technology. One video shows a car that repeatedly and without reason stops in the middle of a street and then drives off again. The explanation came from a passer-by who was carrying a STOP sign sticking out of his bag, misleading the vehicle. In other words, the machine can drive itself, but it lacks the ability to differentiate ‘stop’ signal in different contexts and nuances of incidents on public roads.

 

The loan underwriters at top banks quickly realized that leaving decision making with a compute program is a recipe for disaster as the program has inbuilt biases and no amount of learning can make the program foolproof. Hence in most cases, AI programs are only limited to the first few stages in the process while the critics decision is made by a human.

 

Similarly, Apple credit cards have recently discriminated against women,giving them 50% less credit than men with the same income and profile.

 

Let’s use Netflix as an example. Say you recently watched “Orange is the New Black.” Netflix may then suggest other shows with prison themes, or documentaries about life behind bars, or shows with a strong female lead, etc.

 

Based on past data (your recently watched shows and movies), it’s able to make a prediction about what you will want to watch next. Once you make your latest selection, it will adjust its algorithm to further customize your experience.

 

According to a 2017 scholarly article on Augmented intelligence by Researchers Zheng and Wang, Within augmented intelligence, researchers Define models according to the varied degree of AI and human influence.

 

  • The first classification is called Human-in-the-loop hybrid-augmented intelligence) Human-in-the-loop (HITL) hybrid-augmented intelligence is defined as an intelligent model that requires human interaction. In this type of intelligent system, human is always part of the system and consequently influences the outcome in such a way that human gives further judgment if a low confident result is given by a computer. HITL hybrid-augmented intelligence also readily allows for addressing problems and requirements that may not be easily trained or classified by machine learning.

 

  • The second classification looks at (Cognitive computing based hybrid-augmented intelligence) In general, cognitive computing (CC) based hybrid-augmented intelligence refers to new software and/or hardware that mimics the function of the human brain and improves computer’s capabilities of perception, reasoning, and decision-making. In that sense, CC based hybrid-augmented intelligence is a new framework of computing with the goal of more accurate models of how the human brain/mind senses, reasons, and responds to stimulus, especially how to build causal models, intuitive reasoning models, and associative memories in an intelligent system.

 

Augmented intelligence follows a five-function cadence that allows it to learn with human influence. It repeats a cycle of understanding, interpretation, reasoning, learning, and assurance. Here’s how it works:

  • Understanding: Systems are fed data, which it breaks down and derives meaning from.
  • Interpretation: New data is inputted; the system then reflects on old data to interpret new data sets.
  • Reasoning: The system creates “output” or “results” for new data set.
  • Learn: Humans give feedback on output and the system adjusts accordingly.
  • Assure: Security and compliance are ensured using blockchain or AI technology.

Source: Global Data Magazine

 

It is clear and evident that AI is here to stay. Humans have a deep reliance on Artificial Intelligence which has been around for decades. Yet, the future of Artificial Intelligence is brightest with the humans, morphing in to Augmented Intelligence. Having humans and machines work hand-in-hand is a win-win for both parties.

 

Most scientists and researchers agree that its probably extremely implausible, if not impossible to imagine collective human failure to the extent that AI is allowed to grow unchecked, while simultaneously all other uses of AI beneficial to humankind are ignored.

 

One thing is certain though: despite the ominous predictions and warnings on doomsday scenarios, arising out of the impact of AI on humanity and society, AI will never take over the world or morph in to Terminator style Machines Or Will Smith’s I robot kind of intelligence.

 

]]>
https://1earthtech.com/artificial-augmented-intelligence/feed/ 0 1060
Content Moderation and Social Media https://1earthtech.com/content-moderation-and-social-media/ https://1earthtech.com/content-moderation-and-social-media/#respond Fri, 28 Aug 2020 03:56:26 +0000 https://1earthtech.com/?p=1045 Whats the problem!

Image Credit: Adweek

 

All content platforms and social media companies must keep the content flowing because that is the business model: Content captures attention, provides viewership and generates data (users’ statistics). Content is the starting and the end point of consumers’ journeys on social media. A video, an information post, a tweet, blog post, picture, public service advisories, are all types of content. The platforms then sell that attention (read: viewership), enriched by that data (read: customized ads). But how do you deal with the objectionable, disgusting, pornographic, illegal, or otherwise verboten content uploaded alongside legitimate content?

 

How do Facebook and other tech and social media companies ensure integrity of content on their networks? And how do these companies work to curb misinformation on their platforms about the Coronavirus pandemic or the 2020 elections or any other global or regional event. We have seen state and non-state sponsored actors with nefarious intent take advantage of lax content posting norms.

 

Dangerous fake news has spread on platforms like Facebook in Myanmar, where the Rohingya ethnic minority are persecuted. United Nations has clearly blamed the role of social media in spreading the persecution and this is not the only example of its kind.

 

Misinformation campaigns (aka “fake news”) on Facebook have interfered with democratic elections around the world. After a man used Facebook to live stream his attack on two New Zealand mosques in March 2019, the video quickly spread. YouTube Moderators fought back hard taking down the video as newer versions kept popping up seemingly beating the controls that YouTube has in place to immediately flag already removed material. The uploaders were able to sneak past by using a loophole – exact re-uploads of the video are banned by YouTube, however videos that contain clips of the original footage must be sent to human moderators for review, thereby delaying the process. And again this loophole existed for a purely legitimate reason – to ensure that news videos that use a portion of the video for their segments aren’t removed in the process.

 

In 2017, a live stream on Facebook showed the fatal shooting of a 74 year old retiree in Cleveland, while also showing a man murdering his own child in Thailand. Both videos remained online for hours and racked up hundreds of thousands of views.

 

In a December 2017 report, ProPublica took a revealing look at content moderation. ProPublica gathered from its users 900 examples of where users believed that Facebook content moderation was incorrectly applied. ProPublica then selected 49 of such posts and asked Facebook to explain. Rather shockingly, yet unsurprisingly, Facebook admitted to an error by its moderators in 22 out of 49 posts. Just imagine, 22 out of 49 means approximately 45% or half of all posts in the sample had moderation applied incorrectly. No amount of explaining can explain that.

 

Social media is good, right!

 

Image Credit: Internet

 

Facebook serves as a platform for its billions of regular users to post, view and offer feedback about the content hosted on its servers. But when that content is more “terrorist propaganda” than “brunch photo,” or when it becomes “porn” than “essential context” to an image, the company has struggled to determine the right approach to removing it in time. The traditional methods of company moderators reviewing user-reported infractions is too time consuming, while the AI powered algorithms are too imprecise.

 

With the COVID-19 risk content moderators were sent home, and without proper technology, connectivity, and safety requirements met, Facebook’s automated system took full control. That was an unmitigated disaster, leading to widespread blocking or deleting of posts mentioning Coronavirus from reputable sources such as The Independent and the Dallas Morning News, not to mention millions of individual Facebook users. Those automated systems still have problems.

 

Content from legitimate sources, verified fact-checked sources, and sources with history of posting appropriate and trust worthy content is suddenly being targeted. While there were always instances of some posts getting tagged erroneously, there is an order of magnitude increase in such instances in the post-covid world. Clearly the strategy to have AI and ML based programs call the shots hasn’t worked.

 

“Facebook is blocking COVID-19 posts from fact based sources,” a Facebook source says. On March 17th 2020, according to an Yahoo news article, Facebook suffered from a massive bug in its News Feed spam filter, causing URLs to legitimate websites including Medium, Buzzfeed, and USA Today to be blocked from being shared as posts or comments. The issue blocked shares of some but not all coronavirus-related content, while some unrelated links are allowed through and others are not. Facebook has been trying to fight back against misinformation related to the outbreak, but may have gotten overzealous or experienced a technical error.

 

According to a just released report by NYU Stern, Facebook content moderators review posts, pictures, and videos that have been flagged by AI or reported by users about 3 million times a day. So that is 3 million pieces of content just flagged for review out of possibly billions and billions of content posts. And since CEO Mark Zuckerberg admitted in a white paper that moderators “make the wrong call in more than one out of every 10 cases,” that means 300,000 times a day, mistakes happen.

 

So, is it all an experiment gone wrong? Did the novel coronavirus catch the social media content moderation framework at the worst time?

 

Image Credit: Webhelp

 

The one thing we know for sure is that you can’t control the beast that is Social Media. Generally, the response by firms to incidents and critiques of the social media platforms is primarily ‘We’re going to put more computational power on it,’ or ‘We’re going to put more human eyeballs on it.’” And that is generally fine. For it attempts to resolve the problem, or at least is seen as an attempt to resolve the problem, with or without adequate results. The focus is not on the results, rather on the proclivity to be seen as doing something.

 

Facebook uses more than 70 external partners and fact-checking firms. According to Facebook, it has over 30,000 people working on safety and security — about half of them are content reviewers working out of 20 offices around the world. Facebook employs almost all of these 15,000 content moderators indirectly, mostly outsourced workers. In similar context, YouTube today employs an expected 10 – 12,000 people to patrol al of Youtube and Google’s content.  Similarly, Twitter employees close to 2,000 people in its content review team.

 

Generally speaking, content management or content review falls in to two main buckets. The first is content moderation, where content moderators, mostly contractors working on behalf of lets say Facebook or Twitter, check the content for violations like nudity, sexual content, racism, hate speech, acts of violence or promoting violence, violating laws and community standards, child pornography, and like. Moderators are responsible for reviewing flagged content, and removing it in accordance with the policies of the social media platform. The second bucket is third party fact checking, where Facebook employs more than 70 third party organizations, primarily, news outlets and prominent individuals to check a particular content as True or False. Based on the result then, any one of the many actions can be taken. Either the content is either left up or demoted, or additional labels are added, or additional constraints are placed including monetary impacts or all of the foregoing in extreme cases.

 

According to content management and comprehensive community standards page on Facebook directly, the efforts to moderate and regulate content have three stages. First, is the Policy development process. The content policy team at Facebook is responsible for developing our Community Standards. We have people in 11 offices around the world, including subject matter experts on issues such as hate speech, child safety and terrorism. Many of us have worked on the issues of expression and safety long before coming to Facebook. Second is Enforcement of policies developed previously through its global content moderator workforce. Facebook uses a combination of artificial intelligence and reports from people to identify posts, pictures or other content that likely violates our Community Standards. These reports are reviewed by our Community Operations team, who work 24/7 in over 40 languages. Facebook’s fact-checking rules dictate that pages can have their reach and advertising limited on the platform if they repeatedly spread information deemed inaccurate by its fact-checking partners. The company operates on a “strike” basis, meaning a page can post inaccurate information and receive a one-strike warning before the platform takes action. Two strikes in 90 days places an account into “repeat offender” status, which can lead to a reduction in distribution of the account’s content and a temporary block on advertising on the platform. And finaly, Facebook launched a review process last year. A news organization or politician can appeal the decision to attach a label to one of its posts. Facebook employees who work with content partners then decide if an appeal is a high-priority issue or PR risk, in which case they log it in an internal task management system as a misinformation “escalation.” Marking something as an “escalation” means that senior leadership is notified so they can review the situation and quickly — often within 24 hours — make a decision about how to proceed.

 

If Facebook’s content moderators have three million posts to moderate each day, that’s 200 per person: 25 each and every hour in an eight-hour shift. That’s under 150 seconds to decide if a post meets or violates community standards.

 

Image Credit: TELUS International

 

According the NYU Stern report, and according to some recent investigations by Buzzfeed and news articles by NBC, Forbes and others, the problem of content reviews – whether its content moderation by moderators or third party fact-checking by independent news organizations and individuals is more structural and institutional in nature. The novel coronavirus just exposed a side of it and perhaps aggravated the outcomes.

 

According to a NBC news article, Facebook has allowed conservative news outlets and personalities to repeatedly spread false information without facing any of the company’s stated penalties, according to leaked materials reviewed by NBC News. According to internal discussions from the last six months, Facebook has relaxed its rules so that conservative pages, including those run by Breitbart, former Fox News personalities Diamond and Silk, the nonprofit media outlet PragerU and the pundit Charlie Kirk, were not penalized for violations of the company’s misinformation policies.

 

The list and descriptions of the escalations, leaked to NBC News, showed that Facebook employees in the misinformation escalations team, with direct oversight from company leadership, deleted strikes during the review process that were issued to some conservative partners for posting misinformation over the last six months. The discussions of the reviews showed that Facebook employees were worried that complaints about Facebook’s fact-checking could go public and fuel allegations that the social network was biased against conservatives.

 

“This supposed goal of this process is to prevent embarrassing false positives against respectable content partners, but the data shows that this is instead being used primarily to shield conservative fake news from the consequences,” said one former employee.

 

In a recent case at Facebook, related to appeals process, a Facebook employee filed a misinformation escalation for PragerU, after a series of fact-checking labels were applied to PragerU posts. A Facebook employee escalated the issue because of “partner sensitivity” and mentioned within that the repeat offender status was “especially worrisome due to PragerU having 500 active ads on our platform,” according to the discussion contained within the task management system and leaked to NBC News. After some back and forth between employees, the fact check label was left on the posts, but the strikes that could have jeopardized the advertising campaign were removed from PragerU’s pages.

 

In another case, a senior engineer at one of the top social media giants collected internal evidence that showed the company was giving preferential treatment to prominent conservative accounts to help them remove fact-checks from their content, according to Buzzfeed. The company responded by removing his post and restricting internal access to the information he cited. A week later the engineer was fired, according to internal posts seen by BuzzFeed News.

 

Many employees at top social media companies like Facebook, Twitter and others have expressed deep anguish on their internal inter-company platforms , amid growing internal concerns about the company’s competence in handling misinformation, and the precautions it is taking to ensure its platform isn’t used to disrupt or mislead ahead of the US presidential election.

 

Third party fact checking also suffers from severe debilitating factors severely limiting its outreach. Scale becomes an issue for the fact checkers as most organizations Facebook contracts work of fact checking to, typically only allocates handful of people to the task of fact-checking. Coupled with an impossible amount of fact-checking requests coming in, that means the people are constantly backlogged.

 

According to Sarah Roberts, a pioneering scholar of content moderation, and an information studies expert at the UCLA, the social media companies handle the activities of content moderation in a fashion that diminishes its importance and obscures how the activities of content moderation work. The idea is simple: make it obscure and muddy the waters, to achieve plausible deniability. Something straight out of the play book of top politicians and business executives – plausible deniability. Content moderation is a mission critical activity, yet most social media companies fulfill it with their most precarious employees mostly by just outsourcing the entire journey of content moderation.

 

Just as companies save significant amount of money by outsourcing transport logistics, janitorial and food services, outsourcing content moderation saves these social media giants tons of money. Just as we pointed out earlier, between Facebook, Twitter and Youtube there are close to 40 – 50,000 content moderators. And the number is only growing. Even at conservative estimate of 40,000 people, that is outsourcing work equivalent to 100% of work performed at 4 medium sized outsourcing services providers.

 

The lack of access and the lack of willingness by social media companies to allow any kind of scrutiny of their moderation practices has made content moderation a kind of black box ops where only few people know what takes place. This is certainly by design and it is no accident that the top social media companies choose the convenience of maintaining plausible deniability and the wait and watch approach while incendiary content burns and lights fire to everything around it.

 

What about human content moderators?

 

Image Credit: tampabay.com

 

According to Guy Rosen, VP of Integrity at Facebook, content moderation is a really arduous job. Numerous people have brought this issue to the fore. Watching countless hours of sadistic, violent, disturbing and purely horrific content day in and day out takes its toll. How do you get those hours of visions and thoughts out of your head when you head home? You cannot. Those sights and sounds stay with you. As a content moderator, its really hard to live a normal life after watching 8 hours of non-stop disturbing content.

 

In the recent past, a former Facebook moderator sued, accusing the platform of psychological harm. Former Microsoft employees sued Microsoft for similar reasons after the alleged trauma from reviewing child porn. In a more recent report, The Verge carried out a scathing review of the job conditions for content moderators at Facebook and the harrowing conditions surrounding the job in general. As one employee interviewed in the report put it: “We were doing something that was darkening our soul — or whatever you call it,” he says. “What else do you do at that point? The one thing that makes us laugh is actually damaging us. I had to watch myself when I was joking around in public. I would accidentally say [offensive] things all the time — and then be like, Oh shit, I’m at the grocery store. I cannot be talking like this.”

 

Accenture which performs content moderation for social media companies, has its employees sign a form that directly acknowledges that reviewing such content may be harmful to mental health and could even lead to PTSD.

 

So, is it purely an evil design at play at Social Media giants when it comes to moderating?

Image Credit: VICE

 

To be fair to all, Facebook and other Social media giants do face somewhat of an uphill battle in their efforts of moderating the content. The moment any post, video, or content gets tagged or labeled as requiring fact-checking or misleading or inappropriate, the authors or posters are quick to raise hell about dictatorship, suppression of free speech and infringement of people’s inalienable right of expression.

 

There is always a debate between balancing free speech versus freedom from cruelty and hatred. Or debate between balancing freedom of expression versus right to speak against bullies. A recent attempt by Twitter to mark certain tweets from the President caused a storm and PR crisis. A similar attempt from Facebook recently drew ire of conservatives and put certain ad revenue under threat.

 

Aside from the morals and ethics, at the heart of the debate is a purely financial question: content attracts viewers. More viewers equals more content and vice versa. Any attempt to reduce content, even the borderline inappropriate content will reduce viewers hence impacts revenue. The business models chosen by Facebook, Twitter and Google favor an unremitting, unrelenting drive to add more users and demonstrate growth to investors. More users and more content means more content to moderate and more nuances, but all of that is secondary, a kind of an afterthought.

 

The debate on the usage of internet and governing content uploads is not new. The debate has been going on for some time now and is just about reaching peak interest levels around the world, with many governments promising action like EU and UK; few governments, like China, in fact taking strong action; and few just watching how the entire debate pans out and what, if any, changes come out as result.

 

The big tech players around the world have realized one thing – it’s a tough tight rope walk to control or govern the internet. If a platform puts in too strict controls, through user-reporting mechanism, AI backed algorithms and human monitors flagging and removing content, it will get labeled as ‘dictatorship’ and against free speech. If a platform puts in too few controls, hosting content freely and with little censorship, its going to get run over by activists from all ends of the spectrum, from left to right. It’s quite like an overflowing pot left simmering for long. The only difference is no one can lift the pot and no matter which way its tilted, boiling hot contents are sure to leave scalding marks.

 

With that, lets take a look at The role of Regulators in the efforts to moderate content.

When Mark Zuckerberg wrote the oped in WAPO in March 2019 asking for government and regulators to step in more aggressively to police the internet, he may have elaborated what many insiders feel regarding governing internet, and specifically what content is uploaded for viewers to view, download and use. Yet, not everything seems above board here as the challenges that Mark Zuckerberg cited so eloquently in his oped are the same challenges that have plagued tech industry for years. What has changed recently that governments are being called in to action, while so far the tech industry has fought tooth and nail for freedom of expression and freedom of speech?

 

As the efforts to govern the internet continue, many who are fighting the battle daily are coming to realize the magnitude of difficulty this seemingly simple question of ‘what content to be allowed’ poses. Lets face it: Internet was never known to be deferential to peoples’ preferences. The advocates of freedom of expression and free speech, often big tech companies themselves, fought for as little government control as possible, decrying every move made by governments or regulators around the world.

 

Technology experts, including big tech companies themselves believe that for Zuckerberg and other big tech companies, “regulation” isn’t an uncouth word anymore. As with changing times, the big tech is now embracing regulations, not because of any newfound respect for regulations but purely as a business measure. From early days when big tech companies projected all regulations as reprehensible and fought any and all regulations tooth and nail, to the current day where they are welcoming regulations, the transformation cannot be more melodramatic.

 

Most of big tech today sees regulations as a set of common rules enforced by governments and regulators that’ll allow them to further cement their dominance of the internet. And if anything goes wrong, they always have the comfort of pointing the finger to the” Regulator” big brother.

 

So what is the Way forward to moderate content and to maintain integrity of platforms

According the NYU Stern report, the solution is straight-forward, and calls for increased investment, focus and commitment. The solution is a multi-pronged approach. The first step of this approach begins with Ending outsourcing: to ensure all content moderators are official Facebook or Twitter or employees of the Social media company, with adequate salaries. Increasing the number of moderators significantly is another, as well as placing content moderation under the dedicated oversight of a senior executive.

 

Facebook or Twitter or other Social media companies should also expand oversight in underserved countries, the report suggests.  In addition, the health and well being of content moderators employees should come first. The company should sponsor research into the mental health impacts of moderating the world’s content. And the company should expand fact-checking to curb the spread of misinformation.

]]>
https://1earthtech.com/content-moderation-and-social-media/feed/ 0 1045
Social Media and managing content! https://1earthtech.com/social-media-moderation-content/ https://1earthtech.com/social-media-moderation-content/#respond Wed, 22 Jan 2020 21:51:46 +0000 https://1earthtech.com/?p=781

This article offers a critical view of how content available on the internet quintessentially shapes public opinions and provides an unique challenge to governments around the world.

Introduction

The Social media today has a very different impact potential than it used to 10 years ago. Today the social media is capable of influencing country-wide elections, swaying public opinions, making or breaking brands and people associated with those names, and raising or flat-lining issues. The power wielded by content is truly mind-boggling when one considers that more than 56% of world’s population is online one way or the other. With growing penetration of smartphones and increase in bandwidth accompanied with decrease in data charges, content sharing is becoming ubiquitous and as easy as texting or talking.

 

Governments, regulators and public experts around the world have expressed concern amidst growing need to monitor content which is distributed largely freely today. Who controls what can be seen publicly has been a matter of much debate which is far from settled, however almost everyone involved favors some level of control on viewership of the media content posted online, or in other words, policing the internet. Most through recognize this policing as not a negative or an impediment, rather as a virtue or a social responsibility. Given the power, both direct and indirect, that content platforms usually possess, online available content has become the new hidden weapon on both sides of the debate.

 

The best way to prevent undesirable content from being seen is to never let it be uploaded in the first place. If content publishing platforms like Facebook, YouTube and several hundred others out there are required to screen every single video or post uploaded on their platform, it would simply lose out on millions of subscribers who want to upload their content immediately. If the platforms employ AI powered algorithms to pull down videos, then it must expect errors and content being pulled down or demoted needlessly.

 

Image Credit: Internet

Is Social media uncontrollable?

YouTube works in 80 different languages, is free at the entry level, is available almost everywhere, is extremely popular in the 18 – 24 year age group around the world, even in countries like P.R. China which have their own version of YouTube, and is increasingly being used to ‘monetize’ content like never before. With such mass viewership, Cisco predicts that videos will make up more than 80% of global internet traffic by 2022. YouTube, the world’s foremost content sharing site, has close to 2 billion logged-in users each month (one doesn’t need to be a logged in user to just view videos). Given its low entry barrier, almost everyone with an internet connection has used YouTube at some time.

 

Social media platforms, especially where content is freely distributed, or nearly freely distributed, have had an unfettered run over the past decade leading to their massive popularity today.

 

“We’re just a platform” is almost an extremely convenient way to avoid taking full responsibility for an increasingly serious set of problems. Yet this is exactly the approach that majority of platforms have adapted when it came to own up the responsibility of content moderation and review.

 

Few years back, Steve Stephens recorded himself murdering an innocent victim and then uploaded the footage to Facebook. The horrific act put Facebook under immense pressure to do something. This incident is not alone as there have been may documented cases of teens posting suicides on the internet.

 

Recently, the ghastly attack on the two churches in Christchurch, New Zealand killing more than 50 people was streamed live on Facebook for the world to see. Millions of people watched in disbelief while the lone gunman went on a rampage gunning down unarmed civilians. The video was streamed for several minutes while YouTube Moderators fought back hard taking down the video as newer versions kept popping up seemingly beating the controls that YouTube has in place to immediately flag already removed material. The up-loaders were able to sneak past by using a loophole – exact re-uploads of the video are banned by YouTube, however videos that contain clips of the original footage must be sent to human moderators for review, thereby delaying the process. And again this loophole existed for a purely legitimate reason – to ensure that news videos that use a portion of the video for their segments aren’t removed in the process.

 

For all live streaming and news events, especially ‘Breaking News’, YouTube’s safety team prefers not to use the system for immediately removing child pornography and terrorism-related content, by fingerprinting the footage using a hash system, and rather depends on a system similar to its copyright tool, Content ID, but not exactly the same. The system takes the much longer route – it first searches re-uploaded versions of the original video for similar metadata and imagery. If it’s an unedited re-upload, it’s removed. If it’s edited, the tool flags it to a team of human moderators, both full-time employees at YouTube and contractors, who determine if the video violates the company’s policies. YouTube considers the removal of newsworthy videos to be just as harmful. YouTube prohibits footage that’s meant to “shock or disgust viewers,” which can include the aftermath of an attack. If it’s used for news purposes, however, YouTube says the footage is allowed but may be age-restricted to protect younger viewers, reports The Verge.

 

Image Source: Internet

 

Rasty Turek, CEO of Pex, a video analytics platform that is also working on a tool to identify re-uploaded or stolen content, told The Verge that the issue is how the product is implemented. Turek, who closely studies YouTube’s Content ID software, points to the fact that it takes 30 seconds for the software to even register whether something is re-uploaded before being handed off for manual review. YouTube’s Content ID tool takes “a couple of minutes, or sometimes even hours, to register the content,” Turek said. That’s normally not a problem for copyright issues, but it poses real problems when applied to urgent situations. A YouTube spokesperson could not tell The Verge if that number was accurate.

 

Further, Turek contends, There is no harm to a society when copyrighted things or leaks aren’t taken down immediately. There is harm to a society here though when live-streaming and breaking news puts harmful content out there. This is precisely the reason why live-streaming is considered a high-risk area by Facebook and YouTube and nearly every other platform that provides live-streaming. People need to have different credentials for live-streaming and people violating rules related to live-streaming, who are sometimes caught using Content ID once the live stream is over, lose their streaming privileges because it’s an area that YouTube can’t police as thoroughly. The teams at YouTube are working on it, according to the company, but it’s one the safety team acknowledges is very difficult. Turek agrees.

 

In another incident, Facebook admitted last November that it did not do enough to prevent its platform from being used to “foment division and incite offline violence” in Myanmar against the Muslim Rohingya minority. That came after the United Nations described the events surrounding the mass exodus of more than 700,000 Rohingya people from Myanmar as a “textbook example of ethnic cleansing.”

 

These incidents are not limited to YouTube and Facebook. In China, Bytedance chief executive Zhang Yiming had to issue a public apology in April 2018 after the company was ordered by the central government to close its popular Neihan Duanzi app for “vulgar content”. The company’s Jinri Toutiao news aggregation app was also ordered to be taken down from various app stores for three weeks. Bytedance pledged to expand its content vetting team from 6,000 to 10,000 staff, and permanently ban creators whose content was “against community values”.

 

Image Credit: Internet

Addressing the elephant in the room!

YouTube reportedly employees more than 10,000 people on it’s content moderation and rule enforcement team which represents a 25 percent increase from just about a year ago. according to YouTube took this decision after a number of controversies surfaced over recent past concerning children’s safety on the platform—for those who watch the content as well as those who appear in videos. Yet as other platforms have also struggled with the delicate task of protecting viewers while also ensuring free speech, the big question next year for YouTube is whether this will work.

 

Similarly, Facebook employees close to 15,000 people while Twitter employees close to 2,500 people to moderate the content on its platforms.

 

Relying purely upon humans to act as content moderators and reviewers is not the most viable option. Further the process of reviewing user-reported content is time consuming in itself. Yet, machine learning is not advanced enough at the moment to completely automate the process, though A.I. researchers at Facebook are working on software that they say will eventually enable computers to do most of the work. Current AI and ML systems are not close to being mature enough to address the complexities of moderation, given that almost all intelligence systems operate at a pattern recognition level, where any change in any one of the hundreds of attributes can throw a false positive, thereby passing the content and throwing the system off track. Advanced use cases such as differentiating a jibe from something more serious whether in a video or written format, in an efficient manner that doesn’t unduly stifle free speech is certainly out of question for many years to come.

 

More traditional methods of reviewers depending on user-reported content and user feedback to promote or demote content is too slow and is often regarded as “too-little, too-late“. It could take crucial hours for user-reported infractions to filter through and reach the reviewers, by which time the damage is already done.

Image Credit: Internet

To overcome these challenges, most platforms are now deploying combination of AI with human moderators to come close to building a viable system. The system begins with AI assisted, algorithms based logic program which does an initial flagging and flags anything where it finds even barely perceptible objectionable content. Once the content is flagged, it is taken down immediately or if not yet online, stopped from proceeding further, and queued up for an in-depth review by human operators. The content is then finally posted back online or taken down permanently based on the decision of human moderation process.

 

To illustrate this use of AI, let’s look at Inke, which is one of China’s largest live-streaming companies with 25 million users. As of the end of last year, almost 400 million people in China had done the equivalent of a Facebook Live and live-streamed their activities on the internet. Most of it is innocuous: showing relatives and friends back home the sights of Paris or showing nobody in particular what they are having for lunch or dinner (Source: South China Morning Post). Inke employs 1,200 mostly fresh-faced college graduates who have seconds to decide whether the two-piece swimwear on their screens breaches rules governing use of the platform. The team is the biggest in Inke, accounting for about 60 per cent of its workforce. The content moderators work to detailed regulations on what is allowed and what has to be removed. AI is employed to handle the grunt work of labelling, rating and sorting content into different risk categories. This classification system then allows the company to devote resources in ascending order of risk. A single reviewer can monitor more low-risk content at one time, say cooking shows, while high-risk content is flagged for closer scrutiny.

 

Yet even that – employing AI assisted algorithmic logic coupled with human moderators – may not be enough as Mark Zuckerberg, Facebook’s CEO and majority shareholder, published a memo on censorship in Dec 2018. “What should be the limits to what people can express?” he asked. “What content should be distributed and what should be blocked? Who should decide these policies and make enforcement decisions?” One idea he aired might be thought of as a Supreme Court of Facebook. “I’ve increasingly come to believe that Facebook should not make so many important decisions about free expression and safety on our own,” Zuckerberg wrote. “In the next year, we’re planning to create a new way for people to appeal content decisions to an independent body, whose decisions would be transparent and binding.”

 

Image Credit: Internet

Bigger elephant in the room

There are two major challenges to building and deploying any content monitoring system. The first challenge is simply of scale. As more than 300 hours of videos are uploaded just on YouTube every single minute, the humans monitoring the videos must be capable to correctly identify inappropriate content very, very quickly, virtually in seconds. Making those removal decisions are not simple yet must be made in seconds. Personal biases and tolerances need to be re-calibrated in favor of policies and law. In addition the biggest challenge humans face monitoring content is to maintain their objectivity while watching hours upon hours of videos. Many content reviewers have reported increase in stress levels and allegations of mental illnesses are not uncommon. A group of moderators sued Microsoft in January alleging that they were suffering from PTSD as a result of watching child abuse and other sadistic acts.

 

The second challenge is bigger and is not going to go anywhere anytime soon. As Mark Zuckerberg mentioned in Dec 2018, this challenge is all about who gets to regulate and to what extent, and more importantly who decides the rules of the game? Is it solely up to each content publisher or platform to define their own rules on what content to be flagged as inappropriate and taken down or demoted? Or is it up to the government or a regulatory body to fix this? Zuckerberg feels its clearly the latter.

 

Experts believe that for now at least human curators must continue to work alongside AI. This of course means that both the human moderators and AI backed algorithms have immense ground to cover before a reasonable degree of accuracy can be achieved. It will take time for the platforms, lets say, Facebook, to train its neural networks to streamline that process, while the human moderators continue to face the daunting task of identifying proverbial needle in the haystack.

 

Image Credit: Internet

Conclusion

Lets remember the main problem once again. The social media platforms are responsible for the content hosted on their platform and this responsibility cannot and must not be avoided in the name of freedom of expression, viewers’ choice, viewers’ promoted content, etc or whatever fancy term once may come up with.

 

When anyone speaks about identifying harmful content and the need to censor such content, it feels like the work should be done by robot moderators, but as reality beckons, AI doesn’t yet grasp context and gray areas. The struggle between free speech and censorship keeps humans in the center in this undesirable yet undeniable role.

 

As a content publisher, content aggregator or content reviewer, the stakes have never been higher to allow the content which is legitimate and in the interest of pubic to run freely, while any content which doesn’t meet these two conditions to be taken down as soon as it is put up before the content has the chance to reach any audience.

]]>
https://1earthtech.com/social-media-moderation-content/feed/ 0 781
Big Tech, content moderation and regulations – Three way road! https://1earthtech.com/big-tech-content-moderation/ https://1earthtech.com/big-tech-content-moderation/#respond Tue, 21 Jan 2020 19:45:30 +0000 https://1earthtech.com/?p=779 Big Tech controls lives yet has difficulty controlling what it hosts!

 

Introduction

When Mark Zuckerberg wrote the oped in WAPO in March 2019 asking for government and regulators to step in more aggressively to police the internet, he may have elaborated what many insiders feel regarding governing internet, and specifically what content is uploaded for viewers to view, download and use. Yet, not everything seems above board here as the challenges that Mark Zuckerberg cited so eloquently in his oped are the same challenges that have plagued tech industry for years. What has changed recently that governments are being called in to action, while so far the tech industry has fought tooth and nail for freedom of expression and freedom of speech?

 

All content platforms and social media companies must keep the content flowing because that is the business model: Content captures attention, provides viewership and generates data (users’ statistics). The platforms then sell that attention (read: viewership), enriched by that data (read: customized ads). But how do you deal with the objectionable, disgusting, pornographic, illegal, or otherwise verboten content uploaded alongside legitimate content?

 

The one thing we know for sure is that you can’t do it all with computing. To examine these issues, Roberts pulled together a first-of-its-kind conference on commercial content moderation last week at UCLA. According to Roberts, “In 2017, the response by firms to incidents and critiques of these platforms is not primarily ‘We’re going to put more computational power on it,’ but ‘We’re going to put more human eyeballs on it.’”

 

The debate on the usage of internet and governing content uploads is not new. The debate has been going on for some time now and is just about reaching peak interest levels around the world, with many governments promising action like EU and UK; few governments, like China, in fact taking strong action; and few just watching how the entire debate pans out and what, if any, changes come out as result.

 

The big tech players around the world have realized one thing – it’s a tough tight rope walk to control or govern the internet. If a platform puts in too strict controls, through user-reporting mechanism, AI backed algorithms and human monitors flagging and removing content, it will get labeled as ‘dictatorship’ and against free speech. If a platform puts in too few controls, hosting content freely and with little censorship, its going to get run over by activists from all ends of the spectrum, from left to right. It’s quite like an overflowing pot left simmering for long. The only difference is no one can lift the pot and no matter which way its tilted, boiling hot contents are sure to leave scalding marks.

 

Image Credit: Internet

Governments should Regulate!

Technology experts, including big tech companies themselves believe that for Zuckerberg and other big tech companies, “regulation” isn’t an uncouth word anymore. As with changing times, the big tech is now embracing regulations, not because of any newfound respect for regulations but purely as a business measure. From early days when big tech companies projected all regulations as reprehensible and fought any and all regulations tooth and nail, to the current day where they are welcoming regulations, the transformation cannot be more melodramatic.

 

Most of big tech today sees regulations as a set of common rules enforced by governments and regulators that’ll allow them to further cement their dominance of the internet. And if anything goes wrong, they always have the comfort of pointing the finger to the” Regulator” big brother.

 

Zuckerberg laid out 4 broad areas for government and regulators to govern more effectively and deeply.

Harmful content

The problem: Facebook serves as a platform for its billions of regular users to post, view and offer feedback about the content hosted on its servers. But when that content is more “terrorist propaganda” than “brunch photo,” or when it becomes “porn” than “essential context” to an image, the company has struggled to determine the right approach to removing it in time. The traditional methods of company moderators reviewing user-reported infractions is too time consuming, while the AI powered algorithms are too imprecise. Zuckerberg said he agrees with lawmakers that Facebook has too much power over what constitutes free speech.

 

Mark’s solution: Develop a more standardized approach by setting up an independent body of reviewers who are non-Facebook employees to decide which removal ban stays and to enforce rules. Think a Supreme Court.

 

Objection: The harmful content can be segregated in to multiple buckets. Facebook breaks down the kind of content it’s using AI to proactively detect into seven categories: Nudity, graphic violence, terrorism, hate speech, spam, fake accounts, and suicide prevention. It is unclear if Mark is encouraging the governments and regulators to look at any one category or all of these? The problem becomes even more interesting as at least 5 out of the 7 categories are effectively policed using AI and ML techniques.

 

  • Nudity and graphic violence: These are two very different types of content but Facebook is using improvements in computer vision to proactively remove both. Facebook’s AI backed algorithms depend on computer vision and a degree of confidence in order to determine whether or not to remove content. If the confidence is high, the content will be automatically removed; if it’s low, the system will call for manual review. The degree of confidence is further boosted by the right choices the system makes. Facebook claims that it removes 97% of violence and graphic content, and 96% of nudity before it is reported.
  • Hate speech: Understanding the context of speech often requires human eyes and human understanding as the context is very, very important. Is something hateful, or is it being shared to condemn hate speech or raise awareness about it? Facebook says it has started using technology to proactively detect something that might violate its policies, starting with certain languages such as English and Portuguese. Safety teams then review the content so what’s OK stays up, for example someone describing hate they encountered to raise awareness of the problem. In a major embarrassment, Facebook was used to spread fear and hate in Myanmar against Rohingya Muslims. Today, Facebook says it has hired more language-specific content reviewers, banned individuals and organizations that have broken rules, and built new technology to make it easier for people to report violating content.
  • Fake accounts: Facebook blocks millions of fake accounts every day when they are created and before they can do any harm. This is incredibly important in fighting spam, fake news, misinformation and bad ads. Recently, Facebook started using artificial intelligence to detect accounts linked to financial scams. An account reaching out to many more other accounts than usual; a large volume of activity that seems automated; and activity that doesn’t seem to originate from the geographic area associated with the account are clear warning signals. However, as a point of failure, Facebook failed to recognize several fake accounts set up by a government agency to lure fake admission seekers coming to US to stay long term without proper authorizations. Facebook did launch a complaint later with DHS about the fake accounts
  • Spam: The vast majority of work fighting spam is done automatically using recognizable patterns of problematic behavior. For example, if an account is posting over and over in quick succession that’s a strong sign something is wrong.
  • Terrorist propaganda: The vast majority of this content is removed automatically, without the need for someone to report it first. Facebook is proud of the way its AI systems have been able to remove most terrorist content, claiming that 99% of ISIS or Al-Qaeda content is removed even before being reported by users.
  • Suicide prevention: As explained above, Facebook proactively identifies posts which might show that people are at risk so that they can get help. Facebook claims it passed on information for thousands of suicide attempts to local first responders.

Image Credit: Internet

 

Protecting elections

The problem: Misinformation campaigns (aka “fake news”) on Facebook have interfered with democratic elections around the world. But as the company tries to provide more transparency, it’s having trouble classifying what should or shouldn’t be considered “political.” To be fair to tech companies, censorship is like walking a double-edged sword.

 

Mark’s solution: Let governments set common standards for verifying political actors. The NYT’s Mike Isaac thinks this could work: “When the next erroneous outburst inevitably occurs, Facebook could point toward the law it was forced to follow.”

 

Objection: Facebook has comprehensively shown it can tackle complicated political situations. In the past, the company has complied with requests from leaders of Vietnam and other countries to censor content critical of those governments. In P.R. China, Facebook reportedly created a censorship tool that suppresses posts for users in certain geographies as a way to potentially work with the government. In many western countries, the company has reportedly used the same technology it uses to identify copyrighted videos to identify and remove ISIS recruitment material. Further, Facebook claims that it has developed special programs to give people more information about the ads they see. These special features are since expanded to Brazil and the UK, and will soon in India.

Image Credit: The Verge

 

Privacy -:

The problem: The Cambridge Analytica scandal revealed that Facebook was playing fast and loose with user data. The fine legal print an user electronically signs has often given platforms incredible leeway over how they store and use users’ data while subjected to little or no regulations over how to use the data. There have been numerous breaches as companies have paid little attention to keeping the users’ data secure. To top it all, tech companies have been known to sell the data or allow usage for shady purposes. Worse of all, there are no effective laws in place to prevent misuse and seek preventive and corective actions. New privacy regulations, like Europe’s GDPR law, have come into effect…but that’s leading to an increasingly fragmented internet.

 

Mark’s solution: A global privacy framework à la GDPR.

 

Objection: Facebook and its affiliates through many data deals with parties having need to access users’ data have shown that commercial interests have often succeeded over the privacy concerns. Global frameworks like what Mark has suggested can take years to develop in best case scenario and yet not be applicable uniformly. Mark realizes that this is a gamble he is more than willing to play as its unclear who’s the regulator for a global policy framework?

Image Credit: The Verge

Data portability -:

The problem: Should you, as an internet user, be able to freely move your personal info from one service to another?

 

Mark’s solution: Yes. (Easier to say when he owns WhatsApp, Instagram, and Messenger.)

 

Objection: Again, a solution that serves Facebook and big tech’s own interest is hardly breaking news!

 

Facebook content moderators in action

Image Credit: Internet

 

Monika Bickert and her entire trust and Safety team at Facebook are focused to make sure that their moderators get it right while reviewing the content. There are 60 people dedicated just to crafting the policies for the company’s 15,000 content moderators. These policies are not available publicly anywhere, as these are sensitive and form the “book” which is used by safety teams and the content moderators to refer to when doing their jobs. These policies and definitions of common terms, often nuanced to the level of slangs used in local languages, are revised every two weeks to keep pace with changing ways of how people interact and speak. These sessions are often referred to as mini-legislative sessions, as different teams across the company — engineering, legal, content reviewers, external partners like nonprofit groups — provide recommendations to Bickert’s team for inclusion in the policy guidebook.

 

Many Facebook insiders like Neil Potts emphasize the similarity between what Facebook is doing and what government does. “We do really share the goals of government in certain ways,” he said. “If the goals of government are to protect their constituents, which are our users and community, I think we do share that. I feel comfortable going to the press with that.”

 

Image Credit: The Verge

 

Heart of the matter

In many ways, the change of heart across global tech landscape has resulted from the difficulties the platforms faced while trying to control and govern the harmful content uploaded on the internet’s content platforms. As the efforts to govern the internet continue, many who are fighting the battle daily are coming to realize the magnitude of difficulty this seemingly simple question of ‘what content to be allowed’ poses. Lets face it: Internet was never known to be deferential to peoples’ preferences. The advocates of freedom of expression and free speech, often big tech companies themselves, fought for as little government control as possible, decrying every move made by governments or regulators around the world.

 

Image Credit: Internet

 

Post the 2016 US presidential elections, when the investigations spread in to how agents posing as fake advertisers used social media platforms as a medium to spread biased propaganda to favor one side, the debate around transparency in social media, controls over content available on the internet and the influence on society converged. Google, Facebook, and Twitter were identified as primary targets of these foreign-state sponsored actors as the three companies sat through multiple congressional hearings last year to discuss how foreign agents used and abused their platforms to manipulate Americans. In the aftermath, Facebook tried desperately and failed to fix its fake news problem while it simultaneously reckoned with group of Russian government–backed agents masquerading as advocacy organizations in order to keep pushing socially divisive political messages. Facebook and Google had to publicly apologize and indirectly admit lax controls after it was revealed that advertisers (read: fake) could use their platforms to target a sub-section of society, or people in any age group or ethnicity, based on racist and bigoted interests—like “threesome rape” and “Jews ruin the world.”

 

Facebook hasn’t always taken content review as seriously as they do now. When Facebook Live launched, the technical tool to review the videos did not show where or which part of a video reported for review tended to generate user flags. So, if a Facebook Live video was reported or flagged for inappropriate content, and it was two hours long, the reviewers had to try to skim through entire two hours of content to figure out where the objectionable material might be. Of course, that’s during the pre-election good-old days.

 

Dangerous fake news has spread on platforms like Facebook in Myanmar, where the Rohingya ethnic minority are persecuted. United Nations has clearly blamed the role of social media in spreading the persecution and this is not the only example of its kind.

 

After a man used Facebook to live stream his attack on two New Zealand mosques in March 2019, the video quickly spread. YouTube Moderators fought back hard taking down the video as newer versions kept popping up seemingly beating the controls that YouTube has in place to immediately flag already removed material. The uploaders were able to sneak past by using a loophole – exact re-uploads of the video are banned by YouTube, however videos that contain clips of the original footage must be sent to human moderators for review, thereby delaying the process. And again this loophole existed for a purely legitimate reason – to ensure that news videos that use a portion of the video for their segments aren’t removed in the process.

 

Image Credit: Internet

 

YouTube today employs an expected 12 – 14,000 people, while Facebook now has over 30,000 people working on safety and security — about half of them are content reviewers working out of 20 offices around the world. Similarly, Twitter employees close to 4,000 people in its content review team. Facebook also announced that approx. 1,000 people will work on election focused advertisements alone to prevent the likelihood of the company getting used by false advertisers again.

 

In addition to the manpower, the promised marvels of cutting edge technology, a.k.a. powerful AI based algorithms written specifically to identify, target, and take down or report the harmful content have been, at best under-performers and at worst, simple waste of money and time. And lets not forget this. Facebook used AI backed technology and had an approx. 5,000 strong team of content reviewers before the 2016 elections. Yet it got itself in to significant trouble and scrutiny. Facebook now claims that it has invested record amounts in the past year in keeping people safe and strengthening its defenses against abuse. Facebook also states that its policies have provided people with far more control over their information and more transparency into policies, operations and ads.

 

Image Credit: Vox

 

Facebook also claims that amid all the efforts to control content, it has also started rolling out a content appeals process for users inquiring about their own content that’s been removed. The announcement for content appeals process came much earlier, almost like building hype around the topic is more important than the topic itself. The company said it will expand on this concept so users can appeal decision on reports filed on other people’s content. In addition, Facebook said it is training AI to detect and reduce the spread of “borderline content,” which it describes as “click-bait and misinformation.”

 

The content appeals process announced in 2018 was not a voluntary act of kindness at Facebook, either. Many believe that this is a dual pronged strategy as Facebook executives try to recover their public image. Firstly, this is in response to the increased global scrutiny of how content is treated, sanctioned or legitimized by content platforms. This is also in part due to the lens Facebook has come under especially since The Guardian published a leaked copy of Facebook’s content moderation guidelines, which describe the company’s policies for determining whether posts should be removed from the service. Secondly, Facebook executives hope that such measures would be seen as increasing transparency and fairness, while the opposite is actually true. Moreover, any damage resulting out of such transparent practices will pale in comparison to the risks Facebook will avert.

 

According to Guy Rosen, VP of Product Management at Facebook, “Artificial intelligence is very promising but we are still years away from it being effective for all kinds of bad content because context is so important. That’s why we have people still reviewing reports.” At Facebook, Guy continues, “we continue to train our software system by analyzing specific examples of bad content that have been reported and removed to identify patterns of behavior. These patterns can then be used to teach our software to proactively find other, similar problems.” However these revamped efforts came after multiple reports of algorithms which failed worldwide to identify news with harmful content, instead promoting content inconsistent with the platform’s terms of service, like a man performing sexual acts on a chicken sandwich.

 

Lastly, a look at the role of much touted content moderator has thrown up unmitigated and unexpected, yet serious concerns. What good is a job if it takes away from the professional much more than what it compensates for. What good is a job if it adds more to the healthcare costs and leads to destruction of professional and personal lives? The job of content reviewer itself has come under criticism from many angles. In the recent past, a former Facebook moderator sued, accusing the platform of psychological harm. Former Microsoft employees sued Microsoft for similar reasons after the alleged trauma from reviewing child porn. In a more recent report, The Verge carried out a scathing review of the job conditions for content moderators at Facebook and the harrowing conditions surrounding the job in general. As one employee interviewed in the report put it: “We were doing something that was darkening our soul — or whatever you call it,” he says. “What else do you do at that point? The one thing that makes us laugh is actually damaging us. I had to watch myself when I was joking around in public. I would accidentally say [offensive] things all the time — and then be like, Oh shit, I’m at the grocery storeI cannot be talking like this.

 

Image Credit: Internet

 

An NPR report sheds light on the activities of content moderators as 15,000 moderators contracted by the company worldwide spend their workday wading through racism, conspiracy theories and violence. In the wake of the original story carried out by the Verge, the attention has shifted to work conditions for content moderators. However this attention is not expected to bring about any big results as there are little alternatives to the practice of hiring content moderators to sift through mountains of content daily. As Newton cites in his story, that number is just short of half of the 30,000-plus employees Facebook hired by the end of 2018 to work on safety and security.

 

In a related CNN report, “It’s not really clear what the ideal circumstance would be for a human being to do this work,” said Sarah T. Roberts, an assistant professor of information studies at UCLA, who has been sounding the alarm about the work and conditions of content moderators for years.

 

Facebook would rather talk about its advancements in artificial intelligence, and dangle the prospect that its reliance on human moderators will decline over time. However that’s not going to happen anytime soon as the AI based system isn’t fully ready or dependable yet, hence the pressure is back on humans at least for the foreseeable future.

Image Credit: Internet

 

Conclusion

The roles of government and regulators and the big tech companies have often overlapped. In today’s context however, given the amount of data users are generating, the roles have blurred to some extent. However critical differences have always existed.

 

The big tech companies need to come out of their over-leveraged positions and take due responsibility for the content they host, support and spread. It is not going to be the end of the world if content platforms become more aggressive, nor is it a matter of financial survival as the tech giants have more cash-in-hand than many smaller countries.

 

So, the standoff between AI, human moderators, content platforms hiring the first two, and the target itself – harmful content, continues in to the future. Merely throwing money at the problem can be effective to silence critics in short-term, however long-term improvement plans warrant much more. For one though, the Big Tech has been short of making long term commitments and they have been consistent

 

Image Credit: Internet

]]>
https://1earthtech.com/big-tech-content-moderation/feed/ 0 779
A look at how the Big Tech is building their own Infrastructure! https://1earthtech.com/tech-cable-infrastructure/ https://1earthtech.com/tech-cable-infrastructure/#respond Thu, 02 Jan 2020 02:36:26 +0000 https://1earthtech.com/?p=717 Big Tech companies and investment in Cable Network Infrastructure

 

Introduction

According to some the vast expanse of Internet infrastructure is really just a spaghetti-work of really long wires spread everywhere. While most of the humans now largely experience the internet through Wi-Fi and phone data, the connectivity itself is provided by systems carrying the signals across the world. The signals are transmitted under the ground, carried overhead or travel through deepest oceans. What we can physically see however is just a small part of this mind-numbingly massive infrastructure, for the largest part of internet cabling is virtually passing through the deepest waters.

 

Since the 1990s, the global submarine or undersea cable networks have become the major foundation of worldwide internet traffic and movement of information digitally. The first submarine communication cables laid in the 1850s carried telegraphy traffic, followed by telephone traffic, then data communications traffic. In 1854, installation began on the first transatlantic telegraph cable, which connected Newfoundland and Ireland. Four years later the first transmission was sent – it took nearly 16 hours for the first trans-Atlantic cable sent from Queen Victoria to commemorate the occasion to reach President James Buchanan. Fiber optic cables and communications satellites were both developed in the 1960s, and throughout the cold-war era the undersea cable networks were used by countries to aid and strengthen their communication systems. Internet signals can be carried over satellites in space orbiting around the Earth. There are thousands of satellites in orbit around the Earth today, and the number is increasing. Though fiber optic cables and communications satellites were both developed in the 1960s, and reformed over the years, satellites communications have never been able to get rid of its inherent two-fold problem: latency issues and bit loss in data transmissions. Fiber-optic networks work by sending light over thin strands of glass. Fiber-optic cables, which are about the diameter of a garden hose, enclose multiple pairs of these fibers. Meanwhile, the optical fiber cables can transmit information at 99.7 percent the speed of light. The problems with satellite connectivity, and the advantages of undersea fiberoptic networks have tilted the tide in favor of undersea cables decidedly.

 

In the decades since 1960s, new wireless and satellite technologies have been invented, yet cables remain the fastest, most efficient and least expensive way to send information across the globe.

 

Image Source: Internet World Stats

 

Why is Big Tech interesting in owning the cables

In 2013, Internet traffic was 5 gigabytes per capita; this number is expected to reach 14 gigabytes per capita by 2018. Lets allow that to sink in for a moment.

 

In the modern internet era, where speed, transmission capabilities and robustness of the network held the key, telecom companies did most of the work by building the massive submarine information highways by laying most of the cable under the world’s oceans. During the past decade, however, tech giants (Google, Facebook, Amazon, Microsoft, etc) have started to take more interest in this space by exercising their financial muscle. Experts say submarine cable projects cost up to $350 million, depending on the length of the cable and in case of long projects like the one Facebook just launched, the costs can escalate further quite quickly.

 

According to this NY times article, nearly 750,000 miles of undersea cable already connects the continents to support global insatiable demand for communication and entertainment. Companies have typically pooled their resources to collaborate on undersea cable projects, like a freeway for them all to share.  Alan Mauldin of the research firm Telegeography says only about 30 percent of the potential capacity of major undersea cable routes is currently in use—and more than 60 new cables are planned to enter service by 2021.

 

Image Source: NY Times

 

Traditionally, the relationship between tech companies and telecom provides has been fairly simple.  Like Google’s employees using T-Mobile’s wireless services, the telecom companies built a huge undersea network sinking in billions of dollars, for the bandwidth to be consumed by tech companies to transmit data across the globe. Experts say that as undersea cable technologies improve, it’s not crazy for companies to build newer, faster routes between continents, even with so much fiber already laying idle in the ocean. This model is being disrupted as today, the current growth in new cables is driven less by telecom operators who typically lease the connectivity, and more by companies like Google, Facebook, and Microsoft who are moving from leasing to owning their own connectivity. Tech giants like Google, Facebook, Amazon and Microsoft have always craved ever more bandwidth for multitude of uses – data storage, streaming videos, photos, and other data scuttling between their global data centers.

 

Google is going its own way, in a first-of-its-kind project connecting the United States to Chile, home to the company’s largest data center in Latin America. Alphabet is also reportedly in talks to build its own cable system down Africa’s western coast. Just in 2019 alone, Google is planning to build three underwater cables to help expand its cloud business to new regions. In the past, Google has backed at least 14 cables globally. The company has 13 data centers open around the world, with eight more under construction — all needed to power the trillions of Google searches made each year and the more than 400 hours of video uploaded to YouTube each minute.

 

Image Source: Internet

Ben Treynor Sloss, vice president of Google’s cloud platform, said in a blog post that “together, these investments further improve our network — the world’s largest — which by some accounts delivers 25 percent of worldwide internet traffic,”.

 

Related to the worldwide telecommunications boom and the easy access to 4G networks along with smartphones, more people outside Europe and North America are accessing internet through smartphones in specific and through other means in general. That has prompted companies to think about new growth routes, like between North and South America, or between Europe and Africa, says Mike Hollands, an executive at European data center company Interxion. The Marea cable ticks both of those boxes, giving Facebook and Microsoft faster routes to North Africa and the Middle East, while also creating an alternate path to Europe in case one or more of the traditional routes were disrupted by something like an earthquake.

 

Amazon, Facebook and Microsoft have invested in others, connecting data centers in North America, South America, Asia, Europe and Africa, according to TeleGeography, a research firm.

 

According to the WSJ, in a project dubbed Simba, Facebook is reportedly developing an underwater data cable to encircle the African continent. Along with its newest project, Facebook already has existing undersea cable projects linking North American, European, and East Asian markets.

 

Huawei is launching subsea cable links to Africa.

 

Image Source: Submarinecablemap

 

Benefits keep on stacking up

Experts have pointed out that simply owning its own cables has multiple benefits as having more cables means there are alternate routes for data if a cable breaks or malfunctions. This simple fact has led to a massive re-alignment of priorities among tech giants such as Google, Amazon, Facebook and Microsoft. The new desire to own undersea networks coupled with the massive wallet size and huge drive to take on risks, means that there are less impediments to stop these tech giants from seizing up as much bandwidth as they possible can.

 

Image Source: telegeography

 

In addition to owning the number of undersea cables, the tech giants are investing in the underlying technology itself to improve the speed of data transmissions. Most standard long-distance undersea cables contain six or eight fiber-optic pairs. Google’s new cable dubbed Dunant, is expected to be the first to include 12 pairs, thanks to new technology developed by Google and SubCom, which designs, manufactures, and deploys undersea cables. Google predicts that its cable will transmit around 250 terabits per second which is more than 50% faster than the Facebook and Microsoft’s Marea cable, which transmits data at about 160 terabits per second between Virginia and Spain. Japanese tech giant NEC announced that it has built the technology that will enable long-distance undersea cables with 16 fiber-optic pairs, while Vijay Vusirikala, head of network architecture and optical engineering at Google, says the company is already contemplating 24-pair cables.

 

Not only technologically, even financially it may make more sense to own undersea cables. The companies can increase the amount of data that each fiber pair within a fiber-optic cable can carry while also packing more fiber pairs in to a given cable. As companies pump more and more data through these cables at ever increasing speeds, the value of each cable increases while the cost for each unit of data comes down.

 

Some challenges

It generally takes about a year of planning to chart a cable route that avoids underwater hazards, but the cables still have to withstand heavy currents, rock slides, earthquakes and interference from fishing trawlers. Each cable is expected to last up to 25 years.

 

In total, the volume of undersea cables is hundreds of thousands of miles long; while the cables are laid down at depths as deep as Everest Is tall. The job of laying undersea cables require its own specialized crews, machines and boats. While not the most difficult job, it’s certainly more complex than a matter of dropping wires with anvils attached. The cables must be dropped precisely, generally across flat surfaces of the ocean floor, which can be a daunting task even for the most experienced crews. Care must be taken to avoid anything that can physically disrupt or heavens forbid, damage the cable. Coral reefs, sunken ships, fish beds, and other ecological habitats and general obstructions are just few of such blockers. The cables need to be laid anywhere from shallow sea bed to deepest parts of the ocean where weather conditions can become nasty quickly.

 

The diameter of a shallow water cable is about the same as a soda can, while deep water cables are much thinner—about the size of a Magic Marker. Just like ships’ anchors and fishing trawls, earthquakes can cause significant damage to undersea fiber-optic cables. The damage could be a malfunction or break near or many miles below the surface of the water. When this happens, the telecom operator responsible for maintenance has to find the location of the accident and then work on repairing the impacted sections of the cable depending on the depth where the impacted portion is located. If the cable is in deep waters (6500 feet or greater), the ships lower specially designed grapnels that grab onto the cable and hoist it up for repairs, while If the cable is located in shallow waters, robots are deployed to grab the cable and haul it to the surface.

 

Conclusion

Demand for undersea cables will only grow as more businesses rely on cloud computing services. And technology expected around the corner, like more powerful artificial intelligence and driverless cars, will all require fast data speeds as well. Areas that didn’t have internet are now getting access, with the United Nations reporting that for the first time more than half the global population is now online.

 

]]>
https://1earthtech.com/tech-cable-infrastructure/feed/ 0 717
Removing Bias in Artificial Intelligence: Mission Impossible! https://1earthtech.com/bias-in-artificial-intelligence/ https://1earthtech.com/bias-in-artificial-intelligence/#respond Wed, 13 Nov 2019 22:29:54 +0000 https://1earthtech.com/?p=419 The impact of Bias in Artificial Intelligence!

Originally Published April 21, 2018

 

Artificial Intelligence is broadly referred to as any source, channel, device or usage application whereby tasks normally attributed to human intelligence, such as reasoning, interpretation, storage and processing of information, ability to analyze historical information and ultimately, decision making is reproduced outside human body and human networks.

 

Classic machine learning algorithms involve techniques such as decision trees and association rule learning, including market basket analysis (ie, customers who bought Y also bought Z). Deep learning, a subset of machine learning that includes neural networks, attempts to model brain architecture through the use of multiple, overlaying models.

 

Classic examples of Machine learning include Virtual Personal Assistants like Siri, Alexa, Google, Predictions while Commuting, Videos Surveillance, Social Media Services from personalizing your news feed to better ads targeting, Face recognition, Email Spam and Malware Filtering, Online Customer Support including chat bots, Search Engine Result Refining, Product Recommendations based on previous purchase trends, and Online Fraud Detection including efforts to curb money laundering.

 

From a technology and computer science perspective, bias may refer to the productive bias that enables Machine Learning, both at the level of selecting the training data-set and at the level of training the algorithms. It reminds one of David Wolpert’s ‘no free lunch theorem’. This relates to the trade-off between the size of a training data-set, its relevance, the types of algorithms used, and the accuracy and/or speed of the results. Machine learning research designs involve a number of tradeoffs between e.g. speed, predictive accuracy, over-fitting (low utility) or overgeneralizing (blind
spots), confirming that each choice amongst competing strategies has a cost: there is no free lunch as to the research design for machine learning. From a societal perspective, bias may refer to unfair treatment or even unlawful discrimination. It is crucial to distinguish inherent computational bias from the unwarranted impact of unfair or wrongful bias, while teasing out where they meet and how they interact. This includes an inquiry into the ethical assessments of ML bias, based on the fact that ML applications are re-configuring the ‘choice architectures’ of our online and offline environments. The blind application of machine learning runs the risk of amplifying biases present in data.

 

The underlying assumption of any machine learning program is the existence of an ideal target function or a perfect equation that determines the relationship between input data (e.g., the position of the pieces or board state) and output data (winning or losing the game). In respect to a game of chess, this assumption may hold true, but once we move from chess to human behaviors that are not constrained by a set of unambiguous rules this assumption of a perfect equation is simply wrong.

 

The word ‘bias’ has an established normative meaning in legal language, where it refers to ‘judgement based on preconceived notions or prejudices, as opposed to the impartial evaluation of facts’. The world around us is often described as biased in this sense, and since most machine learning techniques simply mimic large amounts of observations of the world, it should come as no surprise that the resulting systems also express the same bias

 

Human biases are well-documented, from implicit association tests that demonstrate biases we may not even be aware of, to field experiments that demonstrate how much these biases can affect outcomes.

 

Removing bias from AI is the result of deliberate, calculated and thought-out human endeavors, and certainly not an unintended byproduct of certain data analysis. Companies employing any or all five forms of AI — computer vision, natural language, virtual assistants, robotic process automation, and advanced machine learning — must realize that any output they hope to derive are only as good as the data on which the applications are trained.

 

Picture Credit: Google Images

 

Technical breakthroughs and demand for turnkey solutions has led developers to build and deploy platforms where machine-learning engines are made readily available with little or almost no investment in expensive programming teams. AWS or Amazon Web Services recently launched a “machine learning in a box” offering called SageMaker, which non-technical folk can leverage to build sophisticated machine-learning models, and Microsoft Azure’s machine-learning platform, Machine Learning Studio, doesn’t require extensive coding.

 

The intelligent, self-driving systems which rely on complicated algorithms to produce outcomes are as susceptible to the biases as the humans themselves. Just as the human brain builds a model of the world based on what information is fed to it, the algorithms behind machine learning systems and artificial intelligence as a whole build a model world, their own version of reality, based on the data fed to it. If a system is trained on one set of data which is overloaded with samples from a given dataset, the system will have a hard time recognizing other sets of equally valid data. Even in cases where the system is fed large amount of data of several different types, the problem persists as the data may not be deep enough for the system to make unbiased decisions.

 

In a small example of this bias, Google’s photo app, which can apply automatic labels to pictures in digital photo albums, classified images of black people as gorillas. Similarly, Nikon’s camera software misread images of Asian people as blinking. Of course these errors are not intentional, nor are they serious enough to give rise to widespread concerns, yet the bias of these systems has led people to question, perhaps legitimately so, the blind faith many place on artificial intelligence.

 

Picture Credit: Google Images

 

The models learn precisely what they are taught

Considering Bias while feeding data in to machine learning applications is almost becoming a pre-requisite to deploying a machine learning application and is not considered an optional refinement any longer. While machine learning systems enable efficiencies and offer advantages of breakthrough processes, there are many ways in which machines can be taught to do something immoral, unethical, or just plain wrong.

 

To detect the biases in Machine Learning, it is essential to understand how Machine learning actually works, and to detect the series of design choices that inform the accuracy of the outcome. Its essential to understand that each of the design choices that went in to framing the Machine Learning algorithm, entails real life trade-offs that determine the relevance, validity and reliability of the algorithm’s accuracy for real life problems.

 

In one of the early examples of algorithmic bias, 60 women and ethnic minorities were denied entry to St. George’s Hospital Medical School per year from 1982 to 1986, because of a new computer-guidance assessment system that denied entry to women and men with “foreign-sounding names” based on historical trends in admissions.

 

Machine learning in health care holds great promise as it means the avoidance of biases in diagnosis and treatment thereby improving not only the availability of care but the actual results. Health care providers and Practitioners may have bias in their diagnostic or therapeutic decision making. This human bias may be circumvented if a computer algorithm could objectively synthesize and interpret the data in the medical record and offer clinical decision support to aid or guide diagnosis and treatment. In Healthcare, the integration of machine learning with clinical decision support tools, such as computerized alerts or diagnostic support, may offer  targeted and timely information that can improve clinical decisions to health care providers, doctors and nurses. Machine learning algorithms, however, as it is time and often seen are subject to biases. These biases include those related to missing data and patients not identified by algorithms, sample size and underestimation, and mis-classification and measurement error. Given many such examples, there is a growing concern that biases and deficiencies in the data used by machine learning algorithms may contribute to socioeconomic disparities in health care.

 

Types of Biases

 

Lets look at some of the basic types of biases primarily only related to datasets used to train machine learning algorithms.

 

Anchoring bias occurs when choices on metrics and data are based on personal experience or preference for a specific set of data. By “anchoring” to this preference, models are built on the preferred set, which could be incomplete or even contain incorrect data leading to invalid results. Because this is the “preferred” standard, realizing the outcome is invalid or contradictory and can be hard to discover.

Availability bias, similar to anchoring, is when the data set contains information based on what the modeler’s most aware of. For example, if the facility collecting the data specializes in a particular demographic or co-morbidity, the data set will be heavily weighted towards that information. If this set is then applied elsewhere, the generated model may recommend incorrect procedures or ignore possible outcomes because of the limited availability of the original data source.

Confirmation bias leads to the tendency to choose source data or model results that align with currently held beliefs or hypotheses. The generated results and output of the model can also strengthen the confirmation bias of the end-user, leading to bad outcomes.

Stability bias is driven by the belief that large changes typically do not occur, so non-conforming results are ignored, thrown out or re-modeled to conform back to the expected behavior. Even if we are feeding our models good data, the results may not align with our beliefs. It can be easy to ignore the real results.

 

Machines are generally held to be more trustworthy than humans. While a currency counting machine is definitely more accurate and faster than a person counting currency with bare hands, it is difficult to extrapolate the same level of trust to machines with active learning and thinking functions.

 

A look at Allegheny Family Screening Tool: unfairly biased, but well-designed and mitigated

In this final example, we discuss a model built from unfairly discriminatory data, but the unwanted bias may be mitigated in several ways. The Allegheny Family Screening Tool is a model designed to assist humans in deciding whether a child should be removed from their family because of abusive circumstances. The tool was designed openly and transparently with public forums and opportunities to find flaws and inequities in the software.

 

The unwanted bias in the model stems from a public dataset that reflects broader societal prejudices. Middle- and upper-class families have a higher ability to “hide” abuse by using private health providers. Referrals to Allegheny County occur over three times as often for African-American and biracial families than white families. Commentators like Virginia Eubanks and Ellen Broad have claimed that data issues like these can only be fixed if society is fixed, a task beyond any single engineer.

 

Machine Learning and neuron networks are interchangeably used within Artificial Intelligence ecosystems as these represent the way AI is progressing to meet the demands from technology savvy businesses and individuals. Often, these networks create what are known as black boxes, referring to closely guarded virtual boundaries where the internal branching, logic application and evolution of the application itself is obfuscated from public view. And as is often seen, many of the AI algorithms involve use of facial recognition technology, which is deeply flawed in itself.  Facial recognition technology (FRT) which aims to analyze video, photos, thermal captures, or other imaging inputs to identify or verify a unique individual is increasingly infiltrating our lives. Facial recognition systems are being provided to airports, schools, hospitals, stadiums, shops, and can readily be applied to existing cameras systems installed in public and private spaces.

 

 

Additionally, there are already documented cases of the use of FRT by government entities that breach the civil liberties of civilians through invasive surveillance and targeting. Facial recognition systems can power mass face surveillance for the government – and already there are documented excesses, such as explicit minority profiling in China and undue police harassment in the UK. Source: Facial Recognition Technology (Part 1):
Its Impact on our Civil Rights and Liberties, report issued by United States House Committee on Oversight and Government Reform dated May 22, 2019

 

 

Reuters reported a case where a New Zealand man of Asian descent had his photo rejected by an online passport photo checker run by New Zealand’s department of internal affairs. The facial recognition systems registered his eyes as being closed by mistake. When government agencies attempt to integrate facial recognition into verification processes phenotypic and demographic bias can lead to a denial of services that the government has an obligation to make accessible to all constituents. Source: Facial Recognition Technology (Part 1):
Its Impact on our Civil Rights and Liberties, report issued by United States House Committee on Oversight and Government Reform dated May 22, 2019

 

In an investigative study published by ProPublica, the investigators found glaring gaps in how a widely used software that assessed the risk of recidivism in criminals was twice as likely to mistakenly flag black defendants as being at a higher risk of committing future crimes. In other words, simply based on the color of the skin, a machine learnt to classify someone as high risk, precisely similar to what some humans would do.

 

Predictive programs like the one quoted above have generally a higher chance to predicting an outcome on the basis of purely historical data biases leading to re-enforcement of those same biases. In yet another example, in many cities including New York, Los Angeles, Chicago and Miami, law enforcement agencies are using software analyses of large sets of historical crime data to forecast where crime hot spots are most likely to emerge. These areas are then policed heavily in what many argue to be perpetuation of an already vicious cycle of policing over-policed areas to detect and detain more criminals, while crimes in other parts of the city, predominantly white neighborhoods, go relatively undetected.

 

Picture Credit: Google Images

 

AI has been widely used to assess standardized testing in the United States and recent studies suggest that it could yield unfavorable results for certain demographic groups. AI also plays deciding role in hiring decisions, with up to 72% of resumes in the US never being viewed by a human.

 

As recently as 2017, data from the Home Mortgage Disclosure Act showed that applicants from African-Americans are three times as likely and applicants from Hispanic descent are two times more likely to get rejects for conventional loans.

 

 

In more examples, Amazon’s much touted same day delivery service was made unavailable to residents in certain communities based on similar biases in the past, while women job searchers were less likely to see higher paid job ads in results to their job search queries on Google’s search engine than their male counterparts. Both these examples indicate an undesirable outcome for the target audience – though it remains a matter of some speculation if these errors are indicators of broad systemic biases or just glitches in how ad results are displayed at least in Google’s case. Regardless of its findings, the studies that found these errors and biases indicate that there are more, unknown biases occurring out there which are yet to be discovered.

 

 

Lets take a look at Fixing AI Biases

To address potential machine-learning bias, the first step is to adapt transparent means to judge what preconceptions could possibly influence decisions based on a given set of data or what biases currently exist in an organization’s processes, and actively hunt for how those biases might manifest themselves in data. Since this can be a delicate issue, many organizations bring in professionals and external experts to analyze their past and current practices.

 

Up until 10 years ago, the problems of bad or biased data-sets leading to unfair and / or incorrect outcomes was not even considered. The focus was on speed and agility – to build and release systems faster.

 

On the one hand, many people seem to believe that machine learning is agnostic, in the sense of being oblivious to human bias or independent of the design choices that determine its performance accuracy. In that sense, however, machine learning is not agnostic. On the other hand, many people seem to believe that undesirable bias in the training data can be remedied in a straightforward way, thus restoring some kind of neutral
training set, resulting in agnostic machine learning.

 

The errors of intelligent systems are hardly noticeable to most people, yet the presence of similar trends based on data in both historical cases and today’s artificial intelligence systems indicates a deeper malice – something which has gone unnoticed and made its way to how the machines think and act. The problem emerges as the algorithms learn, evolving in to newer and uncharted courses or neuron networks, and take the shape of something which was unimaginable to begin with.

 

Picture Credit: Google Images

 

Breaking the ‘black boxes’ refers to revealing and explaining the ways in which machine learning and neuron networks arrive at data outputs. However this is largely inhibited by the lack of transparency by AI applications to make its entire search parameters and logics open to scrutiny which adds to growing complexity and skepticism around the entire ecosystem of AI. A closer inspection of the prevalent issues within AI makes one fact painfully obvious – Governments and public institutions as well as private individuals need to do more, not just appear to be doing so, but actually undertake conscious initiatives to establish accountability in the system. Most importantly, as the enterprises invest in predictive technologies, they must commit to fairness, due process and transparency.

 

Posing the following questions can help researchers check for systematic bias in underlying data:

  1. Presence of particular groups suffering from systematic data error or ignorance?
  2. Have the researchers intentionally or unintentionally ignored any group?
  3. Are all groups represented proportionally broadly, for example considering an attribute like protected feature of race, are all races being identified or merely one or two?
  4. Has the research team done enough contextual work, for example identifying enough features to explain minority groups?
  5. Has the research team used or is likely to use or create features that are tainted?
  6. Has the research team considered stereotyping features?
  7. Are the data models apt for underlined use case?
  8. Is the data model accuracy similar for all groups under study?
  9. Has the research team identified and corrected predictions that are skewed towards certain groups?
  10. Has the research team optimized all required metrics and not just those that suit the business or potential favorable outcome/s?

 

Finally, just an effort to focus on ethics alone will not do when confronting the framing powers of machine bias, highlighting the need to bring the design choice that determine these framing powers under the Rule of Law.

 

By carefully altering the way different data-sets and groups are assigned to protected or sensitive classes, and ensuring these groups have equal predictive values and equality across false positive and false negative rates, the research team can better detect bias in AI.

 

Humans first: Collective enthusiasm for applying computer technology to every aspect of life has resulted in a tremendous amount of poorly or hastily designed systems. People and companies are so eager to do everything digitally — hiring, driving, paying bills, even choosing romantic partners — that they have stopped demanding that technology actually work to advance the needs of people for whom the technology is created. There are fundamental limits to what humans can and should do with technology.

 

Establishing Anchors: Many applications of machine learning actually work with a so-called “ground truth” to anchor the performance metric; to test whether the system gets it right, machine learning will often require a machine-readable indication of what is “right.” The ground truth is, for instance, based on surveys or interviews where people are asked to assess their own position, emotions, or preferences or, alternatively, based on expert opinion such as medical diagnoses made by medical doctors.

 

Accounting for adversarial training in training phase of Machine Learning -: Adversarial examples exploit the way artificial intelligence algorithms work to disrupt the behavior of artificial intelligence algorithms. In the past few years, adversarial machine learning has become an active area of research as the role of AI continues to grow in many of the applications we use.  In adversarial training, the engineers of the machine learning algorithm retrain their models on adversarial examples to make them robust against perturbations in the data. There’s growing concern that vulnerabilities in machine learning systems can be exploited for malicious purposes.

 

Business Rules, Legal, Regulatory and Compliance framework: Refusing credit, flexible pricing or raising an insurance premium may be based on the freedom to contract, but that freedom is not unlimited and consumer law, competition law, financial services law and insurance law may stipulate further restrictions that must be met, potentially requiring a motivation for a refusal or specific types of price differentiation.

 

Rigorous Testing, Analysis and Adjustments: The study of trade-offs is an important element in the journey to reduce biases. Machine learning research designs involve a number of trade-offs between e.g. speed, predictive accuracy, over-fitting (low utility) or overgeneralizing (blind spots), confirming that each choice amongst competing strategies can be leveraged to tweak the outcome. Is speed more important, or accuracy? Is color of skin given higher weight, or anatomical features? Do bodily movements mean anything? The more important applications should be based on confirmatory research that includes inquiry into causality, so as to prevent delusional inferences that are wrongly taken for granted precisely because there is no understanding of the causal dependencies on potentially unknown parameters.

 

A crucial way to test for biases is by stress-testing the system as demonstrated by computer scientist Anupam Datta of Carnegie Mellon University who designed a program to test whether AI showed bias in hiring new employees. Machine learning can be used to pre-select candidates based on score arrived at through considering various criteria such as skills, ability to lift weights, gender and education. This produces a score which indicates how fit the candidate is for the job. In a candidate selection program for removal companies, where ability to lift weights is a favored requirement, hence the source for bias, Datta’s program analyzed how likely is this bias reflected in the score assigned to each application. The program randomly changed the gender and the weight applicants said they could lift in their application, both crucial parameters for the job. If there was no change in the number of women that were pre-selected by the AI for interviews previously, then it is not the changed parameters that determined the hiring process.

 

As history has shown us, there are hidden pitfalls beyond the general biases in AI systems. Defense and intelligence agencies are overwhelmed by the amount of data generated by surveillance and monitoring systems. An individual being tracked can form unmanageable number of related networks comprising machines and other individuals and keeping a track of all possible nodes of communication multiplied by several hundred thousand subjects can and does become overwhelming. In most glaring cases, certain individuals have slipped through the investigative net, have gone on to inflict severe and long lasting damage through acts of terror. Analysts from one of America’s spy agencies and arguably leading global agency, the NSA, are already overwhelmed by the recommendations of old-fashioned pattern-recognition software pressing them to examine certain pieces of information. Many times its just that the system flags the right individual at the right times, however human analysts or the investigators assigned to probe deeper do not trust the data they are being shown by the software. Having clear explanation of why such individual is flagged and the accurate reasons behind the flag would help convince the investigator and provide rationale for action.

 

Trevor Darrell’s AI research group at the University of California, Berkeley, conducted extensive research with software trained to recognize different species of birds in photographs. Instead of merely identifying, say, a Western Grebe, the software also explains the logic behind its choice – why it thinks the image in question shows a Western Grebe is because the bird has a long white neck, a pointy yellow beak and red eyes.

 

Through a combination of art, research, policy guidance and media advocacy, the Algorithmic Justice League is leading a cultural movement towards equitable and accountable AI.

 

At IBM, through research dedicated to Mitigating human bias in AI, the MIT-IBM Watson AI Lab’s efforts are drawing on recent advances in AI and computational cognitive modeling, such as contractual approaches to ethics, to describe principles that people use in decision-making and determine how human minds apply them. The goal is to build machines that apply certain human values and principles in decision-making.

 

The Gender Shades project evaluates the accuracy of AI powered gender classification products through collaboration with Microsoft, Google and IBM.

 

Picture Credit: Google Images

 

Conclusion

The above biases are not alone. Nor do they represent a significant sample. Observers generally unanimously believe that these are just drops in an ocean of biases living on and growing within the systems collectively known as artificial intelligence.

 

The creators of technology have a living relationship with the technology they create. Machine learning and intelligent systems today take off after their inventors in many cases where the cultural and social rules to inclusion of ethical perspectives, the creators of such systems lend more to it than just their competence. In a way, artificial intelligence reflects the values of its creators. Familiar biases, old stereotypes, vision of external factors and an unfettered outlook of the world have known to become real, tangible traits while a system free of such biases and beliefs exists only in theory books.

 

As some view, this is essentially a fight between conflicting interests. On one hand there are people who are largely interested in maximizing the ROI on investment in technology, even as they mistakenly and painfully ignore the perils from such mindless pursuit of short term goals, while on the other hand there are people who are impacted by the outcomes of biased systems. As in any social-economic conflict, there are people who are unaffected by these biases, shielded by such factors which are favorably treated in the machine learning algorithms, hence are not bothered by the outcomes. At the same time, there are people who though unaffected today, realize the pitfalls and perils form such unfettered misplaced prioritization of developing AI and support the introduction of tighter controls around the way Artificial Intelligence is monitored, controlled and regulated. In a lot of ways, this is also a conflict where those who are aware of technology trends and possess insights must act on the behalf of those who are unaware though may be affected by implications of such biases in AI.

 

Automated systems are not inherently neutral. They reflect the priorities, preferences, and prejudices – the coded gaze – of those who have the power to mold artificial intelligence.

 

We all have an interest in creating robust technologies, AI systems, and other platforms that make our tasks easier and efficient. Further, an understanding of how the data is collected, and the purpose for which it is being used, is paramount to understanding where it can fail or be misused.  On the one hand, we may seek to improve AI to limit the very serious consequences of bias and discrimination — for example, a self-driving car that fails to detect certain pedestrian faces or a greater likelihood that people with darker skin are misidentified as criminal suspects by the police.  At the same time, we must continue to call into question whether that use is supported by our values and should therefore be permitted at all. If we focus only on making improvements to data-sets and machine learning and computational processes, we risk creating systems that are technically more accurate but also more capable of doing harm at an unseen, unmitigated levels.

 

]]>
https://1earthtech.com/bias-in-artificial-intelligence/feed/ 0 419
How Big is Big Data? https://1earthtech.com/big-data/ https://1earthtech.com/big-data/#respond Wed, 13 Nov 2019 19:08:01 +0000 https://1earthtech.com/?p=409 How Big is Big Data?

Originally published April 20, 2018

 

Introduction

When an user visits an internet page, or sends an email, or uploads a picture somewhere, or provides his / her phone number to coffee shop, or swipes credit card at a store, or visits a bank to make deposit or withdraw money or make any other transaction, or pays insurance premium, or buys an air ticket, or train ticket, or, or, or . . .. .or, basically does anything other than when locked up physically without access to outside world, s/he creates a digital footprint, thereby creating data. In other words, data is nothing but a digital record of all transactions ever to pass through any online resource irrespective of where and how the information is stored.

Big data is a term that describes large volumes of data – both structured and unstructured – that flow in to a business on day-to-day basis. But it’s not the amount of data that’s important. It’s what organizations do with the data that matters. Big data can be analyzed for insights that lead to better decisions and strategic business moves.

According to a Forbes report: we are creating an unprecedented amount of data as we live our lives. From social media to the digital footprint we leave as we use services like Netflix or Fitbit or connected systems at work. Every second, 900,000 people hit Facebook, 452,000 of us post to Twitter, and 3.5 million of us search for something on Google.  This is happening so rapidly that the amount of data which exists is doubling every two years, and this growth (and the opportunities it provides) is what we call Big Data.

The sheer value of this data means an industry as well as an enthusiastic, non-commercially driven community has grown around Big Data. Whereas just a few years ago only giant corporations would have the resources and expertise to make use of data at this scale, a movement towards “as-a-service” platforms has reduced the need for big spending on infrastructure. This explosion in data is what has made many of today’s other trends possible, and learning to tap into the insights will increase anyone’s prospects in just about any field.

 

Lets take a look at some big numbers

  • The big data analytics market is set to reach > $73 billion by 2023
  • In 2019, the big data market is expected to grow by 20%
  • By 2020, every person will generate 1.7 megabytes in just a second
  • Internet users generate about 2.5 quintillion bytes of data each day
  • In 2019, there are 2.3 billion active Facebook users, and they generate a lot of data
  • 97.2% of organizations are investing in big data and AI
  • Using big data, Netflix saves $1 billion per year on customer retention

Source: techjury

Image Source: Internet

 

So, what is Big Data, again! The Awesome Vs

Simply put, Big data is collection, storage and analysis of data from traditional and digital sources inside and outside your company that represents a source for ongoing discovery and analysis.

 

Image Source: Internet

 

There is a variety of tools used to analyze big data – NoSQL databases, Hadoop, and Spark – to name a few. With the help of big data analytics tools, we can gather different types of data from the most versatile sources – digital media, web services, business apps, machine log data, etc.

 

1.      Volume of Data

This simply refers to collection of data. Organizations collect data from a variety of sources, including business transactions, social media and information from sensor or machine-to-machine data. In the past, storing it would’ve been a problem – but new technologies (such as Hadoop) have eased the burden.

The data flowing in to the company is most often generated within the company, through one of many channels – sales records, customer database, customer feedback, social media channels, marketing lists, email archives, support systems and any data gleaned from monitoring or measuring aspects of your operations.

Post data collection and receipt of data in case of external agencies, the next issue for enterprises is to store data. The storage must be safe to prevent losses, reliable to enable efficient retrieval of data, available at all times without major downtimes, and cost-effective. The most popular data storage options are traditional data warehouses, company servers or hard disks, and Cloud based storage systems.

Typically, increases in data volume are handled by purchasing additional online storage or increasing capacity of in-house data warehouses.  However, each new unit of data doesn’t provide the same cost-benefit. As with most things, law of diminishing marginal utility kicks in.  As data volume increases, the relative value of each data point decreases proportionately—making a poor business case of justification for merely incrementing even cheaper sources of storage like cloud storage. Viable alternates and supplements include:

 

  • data lakes
  • Implementing tiered storage systems (see SIS Delta 860, 19 Apr 2000) that cost effectively balance levels of data utility with data availability using a variety of media.
  • Limiting data collected to that which will be leveraged by current or imminent business processes
  • Limiting certain analytic structures to a percentage of statistically valid sample data.
  • Profiling data sources to identify and subsequently eliminate redundancies
  • Monitoring data usage to determine “cold spots” of unused data that can be eliminated or offloaded to tape (e.g. Ambeo, BEZ Systems, Teleran)
  • Outsourcing data management altogether (e.g. EDS, IBM)

Source: Gartner Blog

 

Most organizations opt for Cloud-based storage, which is at once, cost effective, offers seamless access, more reliable and readily scalable.  Couple all these advantages with ready availability of IT support to set up storage and respond in case of issues, cloud based storage is a clear winner. Having on site cloud deployment, albeit slightly costlier, is another option for enterprises worried with data security. Overall, Cloud storage is a brilliant option for most businesses. It’s flexible, you don’t need physical systems on-site and it reduces your data security burden. It’s also considerably cheaper than investing in expensive dedicated systems and data warehouses which often require add on investments in teams of people to manage the data warehouse solution deployed. In addition, recurring costs such as maintenance of huge amount of hardware and ever increasing real estate costs make this solution prohibitively expensive for most organizations. Couple this with lack of flexibility to scale storage, and almost impossible amount of work needed to move the data warehouse, its difficult to sell this system to anyone apart from financial institutions, defence and Government buyers.  Cloud storage saves costs because users do not have to buy and administer their own infrastructure. It also allows for flexibility to increase and decrease storage capacity on demand, in real time and without any real impact on performance.

Recent times have also seen development of methods of data storage to manage unstructured data where the schema and data requirements are not defined until the data is queried.

 

2.      Variety of Data

Data sources are varied and as interesting today as they ever were. Almost with each passing day there is more and newer data available to look at same problems. Existing sources of data generate more data with each passing day, while continuously there are newer sources getting added. As more and more existing information gets digitized, and newer sources of data come in to the picture, Conventional boundaries of data management keep vanishing. Traditional data types (structured data) include things on a bank statement like date, amount, and time. The schemas for these kind of data have been long defined and continually improved and adapted. These are things that fit neatly in a relational database.

Structured data is augmented by unstructured data, which is where things like Twitter feeds, social media feeds, audio files, MRI images, web pages, web logs are put — anything that can be captured and stored but doesn’t have a meta model (a set of rules to frame a concept or idea) that neatly defines it.

Unstructured data is a fundamental, disruptive concept in data collection. The best way to understand unstructured data is by comparing it to structured data. Think of structured data as any information or data that is well defined in a set of rules or follows a standard convention. For example, money will always be numbers and have at least two decimal points; names are expressed as text; and dates follow a specific pattern.  With unstructured data, on the other hand, there are no rules. A picture, a voice recording, a tweet — they all can be different but express ideas and thoughts based on human understanding. Further their types and requirements may differ, while the sources of such data are far more numerous than sources of structured data. One of the goals of big data is to use technology to take this unstructured data and make sense of it.

Of course all the continuously mixing of data provides its own challenges. The barriers to effective data management have transformed just like the data itself.  Incompatible data formats, non-aligned data structures, and inconsistent data semantics all require dedicated and focused work to make the data harmonious as well as usable as intended.

 

3.      Velocity of Data

Increasingly, businesses have stringent requirements from the time data is generated, to the time actionable insights are delivered to the users. Therefore, data needs to be collected, stored, processed, and analyzed within relatively short windows – ranging from daily to real-time

So now you have the data that’s been collected and stored securely. What to do with it next? Unless you have a magical wand, the data wont translate itself in to usable insights. Rather, it has to follow what is perhaps the most critical step of the value chain – analysis of mountains of data. Most companies do this with the use of data scientists and specialist software from companies such as IBM, Oracle and Google, as well as host of start-up players who offer more agility and customization than their heavyweight counterparts. Most of the data crunching software in the market today are designed around ‘dummies’, i.e. for people with less specialist knowledge and for smaller enterprises who may not want to invest in hiring expensive resources like data specialists and scientists. The purpose of data analysis is to produce insights for the business as well as highlight actions for the business to contemplate. The ROI on the software is tied critically to these two factors as well as the ease of use and overall responsiveness of support teams.

Analytics have long provided valuable business insight, but it’s been too time consuming and expensive for mass adoption in some cases. The process of extracting data, transforming it into something that fits operational needs, loading and conducting analysis can take weeks, sometimes even longer. Due to the requirements of today’s online, on-demand businesses, that’s not an option any more if businesses are to retain their competitive advantage. Data is growing too fast, the potential insights are too valuable and competition is too fierce for IT systems to have multi-week wait times to perform analytics. That’s driving the development of new solutions, such as near real-time analytics based on Hadoop.

Originally, data frameworks such as Hadoop, supported only batch workloads, where large datasets were processed in bulk during a specified time window typically measured in hours if not days. However, as time-to-insight became more important, the “velocity” of big data has fuelled the evolution of new frameworks such as Apache Spark, Apache Kafka, Amazon Kinesis and others, to support real-time data processing.

Data visualization is the process of passing findings, outputs and actionable inputs from the previous process of data analysis to the decision makers or other stakeholders who need to have the information. Normally, this is in the form of presentations, summary reports, brief reports, charts, figures, or any other form as per the organization / department culture.

The key at this stage is to present the key actionable insights to top stakeholders as clearly as possible. Many a great data collection, storage and analysis projects fail to deliver their goal due to lack of clear and concise information or otherwise burying of actionable inputs under heaps of relatively less useful information. It’s clearly unrealistic to expect busy people to wade through mountains of data with endless spreadsheet appendices and extract the key messages. Remember: if the key insights aren’t clearly presented, they won’t result in action.

 

4.      Veracity of Data

According to a survey, 1 in 3 business leaders do not fully trust the information they get to make the decisions they need to make to operate their business more efficiently.

Big Data Veracity refers to the abnormal biases, noise and abnormality in data that can and does materially affect the integrity of data, hence affecting the outputs produced through analysis of such data. Is the data that is being stored, and mined meaningful to the problem being analyzed? In scoping out the big data strategy, the need to have employees and partners work to help keep data clean and processes to keep tainted data from accumulating in systems.

In addition to the above 4 all-important Vs of Big Data, off late many industry professionals have proposed additional characteristics in the form of other Vs like Volatility, Variability, Value, and Validity. Though these are some interesting viewpoints, none of these new Vs represent much definitional departure from existing Vs, meaning that these can be absorbed within the definitions of existing Vs.

 

Image Source: IBM

Getting in to Big Data: Examples and Uses

From traffic patterns, online shopping, music downloads, user’s web history and medical and social records, data is recorded, stored and analyzed to enable the technology and services that the world relies on every day.

It’s only in the past few years that big data has really started to break out. That’s because we finally have computers and connections fast enough to transfer and analyze the mounts of collected information.

Depending on the industry, organization and intended usage, big data encompasses information from multiple internal and external sources such as databases, user preferences, user records, transactions, social media, enterprise content, sensors and mobile devices. Companies can leverage data to adapt their products and services to better meet customer needs, optimize operations and infrastructure, and find new sources of revenue.

 

Image Source: Internet

The English Premier League has upped the ante with data trackers embedded in players’ uniforms that monitor performance in real time. Coaches are using this information to predict everything from long term performance and likelihood of injury to how a player will perform in the next game.

Healthcare is a prime example of an industry benefitting from Big Data breakthroughs. Big Data is propelling incredible advancements in not only preventive but curative medicine as well. There are many applications where deep learning has enabled machines to scan thousands of medical records, cluster the data in many different ways, and use statistics to help medical professionals make valuable inroads in to disease treatment and prevention. The machines use these patterns and probability models to predict whether a patient is likely to be affected by a disease or not. Climate change is yet another good example. We have data on socio-economic populations, historic temperature measurements, flood and drought statistics, natural calamity, and others. These data sets are routinely analysed, updated and compared against different time periods and conditions to accurately predict future weather conditions. Once all that data is connected and analyzed, it is expected to give us many insights that will help us to understand global climate patterns or even how to restore natural resources or fauna.

Structured data already contain huge troves of information on security, transactions and compliance. This data can be readily mined by machines, further processed in quick time with complex algorithms to detect anomalies and fraud and issue safety alerts. A typical example is credit card usage, where an abnormal use or higher than normal amount can trigger a safety alert in real time to the bank or financial institution which normally places a hold on such transaction until it can verify the usage. Further, the data generated by credit card or electronic transactions helps gather operational intelligence on customer preferences.

Just few years back, internet searches were difficult to master. Users had to search using combinations of keywords and shift through multiple online resources to find what they were looking for. Today, thanks to indexing and super-fast searches, the probability of getting accurate results in fraction of seconds is almost 100%.

Industrial uses of Big Data include analyzing the behavior of several machine parts to predict failures, forecast replacements, schedule and synchronize replacements, and automatically order new parts using cloud-based ERP systems. This is how new data-driven systems help eliminate unplanned downtime and streamline supply chain processes.

Social media, has grown almost exclusively on the edge offered through Big Data analytics. The “always on” platforms continuously monitor and record individual connections, preferences, places of interests, priorities etc. The social graph data is stored in different storage tiers (memory, solid-state drives, and magnetic media and others), transformed and then mined with complex mathematical algorithms. The machines then predict individual behavior/decisions in a given environment, which help companies to push personalized advertisements and generate revenue.

Source: Western Digital Powerful ways data is changing our world

 

Conclusion

By 2026, the big data market is expected to top $92 billion. That’s a 307.8% increase over the $22.61 billion the market was worth in 2015.

The healthcare analytics market is projected to reach $24.6 billion by 2021. That’s a 187% increase from the $7.39 billion it was in 2016. And it represents a compound annual growth rate of 27.1%.

The uses of Big Data are virtually unlimited. From smart cities to international financial markets, to medical and pharma research, to agriculture and social innovations, the journey has just begun. In the world of data, significant changes are happening — some are obvious while others are below the surface. We’re only just starting to see how revolutionary big data can be, and as it truly takes off, we can expect even more changes on the horizon.

The ever changing landscape of ‘Digital’ today has led to mind-boggling amounts of data. The application surge, growth of e-commerce, virtually unlimited sources of unstructured data, a rise in merger & acquisition activity, increased collaboration, and the drive to harness information as a competitive catalyst for growth and innovation is driving enterprises to higher levels of consciousness about how data is managed at its most basic level. The effect of the e-commerce surge, a rise in merger & acquisition activity, increased collaboration, and the drive to harness information as a competitive catalyst towards both growth and innovation is driving enterprises to higher levels of consciousness about how data is managed at its most basic level. The top level strategy is often focused at bottom level data usage.

 

 

]]>
https://1earthtech.com/big-data/feed/ 0 409