In many busy households around the world, it’s not uncommon for children to shout directions to Apple’s Siri or Amazon’s Alexa. They can make a game of asking the voice-activated personal assistant (VAPA) what time it is or request a popular song. While this may seem like a simple part of home life, there is much more going on.
VAPAs continuously listen to, record and process acoustic events in a process that has been called “eavesdropping”, a portmanteau of eavesdropping and data collection. This raises serious concerns about privacy and surveillance issues, as well as discrimination, as audio tracks of people’s lives are collected with data and scrutinized by algorithms.
These concerns are amplified when we apply them to children. Their data accumulates over a lifetime in ways that go beyond what was ever collected about their parents, with far-reaching implications we haven’t even begun to understand.
Read also: Amazon wants you to take Alexa on the go now
I always listen
Adoption of VAPA is happening at a staggering pace as it includes mobile phones, smart speakers and the ever-increasing number of products that are connected to the Internet. These include children’s digital toys, home security systems that monitor for break-ins, and smart doorbells that can pick up conversations on the sidewalk.
There are pressing issues arising from the collection, storage and analysis of sound data as it relates to parents, youth and children. Alarms have been raised in the past—in 2014, privacy advocates raised concerns about how much the Amazon Echo was listening, what data was being collected, and how the data would be used by Amazon’s recommendation engines.
Yet despite these concerns, VAPA and other eavesdropping systems have spread exponentially. Recent market research predicts that by 2024, the number of voice-activated devices will grow to over 8.4 billion.
Recording more than just speech
More than just spoken statements are collected, as VAPA and other eavesdropping systems eavesdrop on personal characteristics of voices that inadvertently reveal biometric and behavioral attributes such as age, gender, health, intoxication, and personality.
Information about an acoustic environment (such as a noisy apartment) or specific sound events (such as breaking glass) can also be gathered through “auditory scene analysis” to make judgments about what is happening in that environment.
Wiretapping systems already have recent experience cooperating with law enforcement and obtaining subpoenas for data in criminal investigations. This raises concerns about other forms of creeping surveillance and profiling of children and families.
For example, data from smart speakers can be used to create profiles such as “noisy households,” “disciplinary parenting styles,” or “troubled youth.” In the future this could be used by governments to profile those who rely on social assistance or families in crisis with potentially severe consequences.
There are also new eavesdropping systems touted as a child safety solution called “aggression detectors.” These technologies consist of microphone systems loaded with machine learning software, questionably claiming to be able to help predict incidents of violence by listening for signs of increased volume and emotion in voices, as well as other sounds such as eg breaking glass.
Read also: Amazon has a plan to make Alexa imitate someone’s voice
Monitoring of schools
Aggression detectors are advertised in school safety magazines and at law enforcement conventions. They have been deployed in public spaces, hospitals and high schools under the guise that they can prevent and detect mass shootings and other cases of deadly violence.
But there are serious issues surrounding the efficacy and reliability of these systems. One brand of detector repeatedly misinterpreted children’s vocal cues, including coughing, screaming and clapping, as indicators of aggression. This raises the question of who is protected and who will be made less secure by its design.
Some children and young people will be disproportionately harmed by this form of securitized listening and the interests of all families will not be equally protected or served. A recurring criticism of voice-activated technology is that it reproduces cultural and racial biases by imposing voice norms and misrecognizing culturally different forms of speech in relation to language, accent, dialect, and slang.
We can expect that the speech and voices of racist children and youth will be disproportionately misinterpreted as aggressive sounding. This troubling prediction should come as no surprise, as it follows deep-seated colonial and white supremacist histories that have consistently maintained a “sound color line.”
Sound politics
Eavesmining is a rich site of information and surveillance as the sonic activities of children and families have become valuable sources of data to be collected, monitored, stored, analyzed and sold without the subject’s knowledge to thousands of third parties. These companies are profit-driven, with little ethical obligation to children and their data.
Without a legal requirement to delete this data, the data accumulates over children’s lifetimes, potentially remaining forever. It is not known how long and how extensively these digital traces will follow children as they age, how widely this data will be shared, or how much this data will be cross-referenced with other data. These issues have serious implications for children’s lives both now and as they age.
There are countless threats posed by wiretapping in terms of privacy, surveillance, and discrimination. Individualized recommendations, such as information privacy training and digital literacy training, will be ineffective in addressing these issues and place too much onus on families to develop the necessary literacy to counter eavesdropping in public and private spaces.
We need to consider developing a collective framework that addresses the unique risks and realities of wiretapping. Perhaps the development of principles of honest listening—an auditory spin on the “principles of honest information”—would help to appreciate the platforms and processes that influence the sound lives of children and families.
By Stephen J. Neville and Natalie Coulter, York University, Canada (The Talk)
Add Comment