Data, Bias, and Delegated Judgment

Argument layer

Deep Read

The structure beneath the prose

This route keeps the complete chapter visible while exposing concepts, inferences, objections, and places where your own judgment must do work.

From Data Trails to AI Systems: How Human Activity Becomes Prediction, Automation, and Generation

Rosa opens Canvas on her phone during a break at work. She checks the deadline for a philosophy discussion, downloads the reading, and closes the app before the page finishes loading. Later that night she watches part of a lecture video, pauses it when her child wakes up, and finishes the reading from a PDF she saved to her phone. The next morning she opens the quiz twice before submitting it. She turns in the discussion post at 11:47 p.m. Two days later, she ignores an automated reminder because she has already handled the problem in another class. She also has a financial-aid hold she does not fully understand.

None of these events is dramatic. None is an AI decision. Each one looks like ordinary college life.

But the college and the tools around it may record parts of the week: logins, timestamps, submissions, grades, video views, advising notes, device information, aid status, and messages. Some of Rosa’s work is visible to the system. Some of it disappears. Canvas may record the click on a reading, while the hour she spent reading offline stays invisible. A dashboard may see a late submission while missing the work schedule behind it. A risk model may see a pattern of missed activity while missing the cracked phone screen, the bus transfer, the childcare interruption, the financial-aid confusion, and the fact that Rosa is actually doing the reading.

This is the beginning of datafication. A week of human activity becomes a record. The record becomes searchable, comparable, and reusable. It may help an instructor notice confusion earlier. It may help an adviser reach out before a student disappears. It may help a college see where students get stuck. It may also create a thin profile of Rosa that follows her into a dashboard, a nudge, a prediction, a generated summary, or a decision she never sees.

The ethical question begins before the algorithm. Before we ask whether an AI system is fair, accurate, biased, useful, or harmful, we need to ask how the data path was built. This chapter follows that path from lived activity to trace, record, category, proxy, prediction, generated output, or action.

The Week And The Record

Datafication means turning activity, behavior, relationships, text, images, movements, sounds, bodies, places, or institutional events into data that can be stored, linked, compared, analyzed, predicted from, or used to generate outputs. The word sounds abstract, but the process is familiar. A step count turns movement into a number. A gradebook turns learning into scores. A streaming service turns watching into preference data. A pharmacy coupon turns a medication search into a health-related data point. A navigation app turns a route into location history. A chatbot prompt turns a question into text that may be logged, reviewed, retained, or used under certain product settings.

The simplest path looks like this:

A trace is not yet the whole story. A trace is a small captured sign: a click, timestamp, quiz score, location ping, search query, upload, like, pause, purchase, prompt, sensor reading, or note. A record is a stored trace or collection of traces organized so that someone or something can use it later. A record may be accurate and still incomplete. Rosa really did submit at 11:47 p.m. That does not mean the late timestamp explains her academic situation.

Datafication is useful. It is one reason modern systems can personalize services, coordinate support, improve search, recommend resources, detect patterns, preserve institutional memory, make inaccessible information easier to find, and operate at scale. A college cannot notice course-level bottlenecks without records. A doctor needs durable traces of a patient’s history. A library improves search by learning something about sources, queries, and use. Navigation apps need location and traffic data to suggest faster routes. Large language models depend on large bodies of prior text, images, code, structured data, feedback, and human labeling.

Data and algorithms are powerful because they help systems see patterns no individual person could easily see alone. They can make a service more responsive, a process easier to inspect, or a tool more accessible. They can also make people easier to sort, rank, profile, monitor, ignore, or act on at a distance.

That double edge is the center of this chapter. Datafication is a capability amplifier. It gives institutions, companies, researchers, governments, and AI systems more ways to see and act. That added capacity can support care, discovery, creativity, accountability, access, and coordination. It can also make bad categories, weak proxies, hidden labor, and unequal power easier to automate.

The chapter will keep returning to Rosa’s week because it is ordinary. Most AI ethics problems do not start with a movie-style robot or a spectacular scandal. They start with records. They start when some part of a person’s life becomes legible to a system.

Data Is Made

It is tempting to talk about data as if it were already sitting in the world, waiting to be collected. That picture is too simple. Data is produced. Someone decides what to record, how to format it, which categories to use, how long to keep it, how to combine it, what counts as an error, who can access it, what later use the record can support, and what the data will not try to represent.

Lisa Gitelman’s edited collection “Raw Data” Is an Oxymoron is often cited because its title captures the problem. Data may feel raw when it arrives in a spreadsheet or dashboard, but it already passed through choices. A reading log exists because a platform records certain events. A grade exists because an instructor used a particular assignment, rubric, deadline, and submission system. A risk score exists because someone decided which traces should count as evidence of risk.

That does not make data fake. It means data has a history.

Danah boyd and Kate Crawford make a related point in “Critical Questions for Big Data”. Large datasets can create new knowledge, but size does not remove assumptions. A huge dataset can still be incomplete, biased, poorly interpreted, or badly matched to the question being asked. In student language, more data is not the same as better judgment.

Return to Rosa’s week. The record may show that she watched only part of a lecture video. It may not show that she already understood the material, that she watched with a friend, that she downloaded the transcript, or that she read the chapter offline. The record may show a late submission. It may not show that she submitted late because her work shift changed. The trace may be accurate while still leaving out the context needed for judgment.

Jose van Dijck’s work on datafication, dataism, and dataveillance is useful here because she shows how digital platforms encourage people to treat data as a privileged way of knowing. If a thing can be counted, tracked, graphed, or predicted, it begins to look more real than the parts of life that resist measurement. That can help. It can also distort.

The translation from lived activity into data often makes hidden patterns visible. A college may discover that students in a course consistently stall at the same assignment. A tutoring center may discover that evening hours are more useful for working students. A public-health system may see a pattern of illness earlier than individual clinics would. A creator may learn which parts of a tutorial confuse viewers. Data can reveal bottlenecks and make institutions answerable for patterns they would otherwise miss.

The same translation can flatten a person. A record can preserve one trace while dropping the context that made the trace understandable. A dashboard can help someone notice Rosa and still fail to understand her. Ethical judgment begins with that translation.

Data-making also involves labor. Some labor is visible: a researcher designs a survey, a nurse enters a clinical note, a teacher grades an assignment, a clerk updates a file. Some labor is hidden: users produce behavioral traces by using a platform, moderators remove disturbing content, contractors label images, students generate examples, writers make public work, and technicians clean datasets. When a system later treats the dataset as a resource, the people who made it useful can disappear from view.

Why Institutions Count: A Short Genealogy Of Data Power

Rosa’s Canvas record belongs to a much older story. Data power began long before AI, computers, databases, or apps. Human societies have used records to count people, goods, taxes, land, births, deaths, crimes, illnesses, debts, religious membership, property, movement, communication, and work. The tools changed. The ethical pattern stayed recognizable: records help institutions coordinate life at scale, and the same records can make people easier to govern, extract from, rank, discipline, or ignore.

Ancient states counted because administration requires memory. Census histories often point to Mesopotamia, Egypt, China, Rome, and other early states as places where rulers used counts to manage food, labor, taxes, land, and military obligations. The UK Office for National Statistics describes ancient census-taking in connection with provisioning, taxation, and clay-tablet records. The Population Reference Bureau notes the often-cited Han Dynasty census of 2 CE, which recorded households and population on a massive scale for its time. The Inca used quipu, knotted-cord record systems, to encode numerical information in a decimal structure.

Those examples should not be forced into a single origin story. They show a recurring administrative move. A count can help distribute grain, plan roads or irrigation, assign labor, levy taxes, recruit soldiers, or organize civic status. The same count can support public goods and strengthen control. Every count carries a theory of what the institution wants to see.

Older texts already show unease about counting. The Hebrew Bible includes census scenes where numbering a people is connected with order, military strength, royal power, and divine judgment, including Numbers 1, 2 Samuel 24, and 1 Chronicles 21. Those stories use moral categories different from modern privacy law, but the concern is familiar enough. Counting people can become an act of care, stewardship, pride, distrust, preparation for war, or political control. The act of counting is never only technical once the count changes what an institution can do.

Medieval and early modern recordkeeping moved the same pattern into churches, monasteries, royal offices, courts, estates, and standardized state forms. Parish and church registers recorded baptisms, marriages, deaths, and community membership. A register could preserve family history, support inheritance claims, and make care more durable. It could also make belonging subject to institutional authority. The Domesday Book of 1086 gives a sharper example of recordkeeping as governance. William the Conqueror’s survey recorded land, holders, resources, people, and value across England. The record could settle claims and clarify obligations. It could also make taxation and control more precise.

Early modern states added standardization. Forms, tables, maps, ledgers, printed instructions, and regular censuses made people and places comparable across distance. A local description could be translated into a table. A household could be placed inside a category. A town could be compared to another town. Iceland’s 1703 census and Sweden’s 1749 population tables are often cited as early European examples of systematic population recordkeeping. The United States built a decennial census into the Constitution beginning in 1790; Britain followed with its first modern census in 1801.

James C. Scott’s Seeing Like a State helps explain the ethical tension. States often simplify local life so it can be seen from the center. Standard names, standard measures, cadastral maps, population tables, and official categories make administration possible. They can also flatten local knowledge and make a complicated life easier to act on from far away. Legibility is useful because institutions cannot respond to what they cannot see. It is dangerous when the visible part becomes the whole person.

The nineteenth century added faster communication and more durable identification. Telegraphs and telephones moved messages across distance. They also created new possibilities for interception and metadata: who contacted whom, when, from where, and how often. Photography, police files, anthropometry, and fingerprints turned bodies into records that could be stored, searched, and compared. The National Library of Medicine’s account of Bertillon’s identification system shows how physical measurements and photographs became standardized identity records. The NIJ Fingerprint Sourcebook traces fingerprinting through colonial administration, policing, and forensic identification. Warren and Brandeis’s 1890 article on “The Right to Privacy” responded to a world in which photography and newspapers were making personal life more exposed.

Then recordkeeping became machine-readable. Herman Hollerith’s punch-card tabulators were used for the 1890 U.S. Census. The Census Bureau explains that Hollerith’s machine read holes in paper cards to tabulate census data, including individual characteristics and cross-tabulations. That shift matters because it turned administrative records into inputs for automated processing. A census form became a card. A card became a signal in a machine. A machine could count combinations faster than clerks could. Later, the Census Bureau’s use of UNIVAC I moved census processing into electronic computing.

This is the bridge to modern big data. Digital records are cheap to copy, search, combine, and reuse. Platforms turned ordinary life into streams of observed behavior: searches, clicks, pauses, likes, locations, purchases, uploads, messages, prompts, ratings, and watch histories. Some data is volunteered. Some is observed. Some is inferred. Some is purchased, joined with other records, and sold downstream. The FTC’s 2024 staff report on social media and video-streaming services makes this platform pattern concrete: selected major services collected, retained, shared, and monetized extensive user data, while safeguards were weak in several areas.

Large language models add another layer to the same history. They are built from data trails, but the trails now include cultural material: websites, books, code, images, posts, documentation, public records, licensed corpora, synthetic data, user feedback, and human labeling. The model does not merely store a record in the old sense. It learns patterns from vast collections and can generate new text, code, images, summaries, plans, or classifications. The old recordkeeping question becomes sharper: where did the material come from, what did it leave out, who labored to make it useful, who had authority to reuse it, and what action follows from the generated output?

The genealogy matters because it blocks two lazy stories. One says datafication is simply surveillance and should be rejected wherever it appears. That misses the public goods that records can support: representation, health, safety, access, institutional memory, research, accessibility, planning, and care. The other says datafication is neutral because records are only facts. That misses how records are made, categorized, preserved, combined, inferred from, and used by institutions with unequal power.

The better habit is to trace the data path. Ask what was counted, what category made it usable, what institution gained capacity, what context was dropped, what later use became possible, and who could inspect or challenge the result. Rosa’s Canvas record is new in its medium, but not in its moral structure. A life becomes legible to a system. The system becomes more capable. The question is what that capacity is for, how bounded it is, and whether the people made legible can answer back.

Categories Make People Legible

Institutions need categories. A college needs to know who is enrolled, who has completed a prerequisite, who qualifies for aid, who needs accommodation, who has submitted work, who is graduating, and which courses fill too quickly. Without categories, the institution cannot coordinate services or treat similar cases consistently.

Categories can create transparency. If a college tracks pass rates across course formats, it can see whether an online course is creating barriers. If it tracks advising wait times, it can see whether students are being served. If it tracks withdrawal patterns, it can ask whether a policy, schedule, or course design is producing avoidable harm. Data can make institutional patterns visible enough to challenge.

The same categories can also misread people. A student marked “inactive” may be reading offline. A student marked “at risk” may be dealing with a temporary work conflict. A student marked “nontraditional” may be treated as an exception even when students like her are common at a community college. A student marked “engaged” may have clicked every page and understood very little.

Legibility is the name for this condition: a person, action, place, object, or situation becomes readable to an institution or system. Legibility is useful because institutions cannot respond to what they cannot see. It is dangerous when the institution starts treating the legible part as the whole person.

Geoffrey Bowker and Susan Leigh Star’s Sorting Things Out gives students a way to read categories as infrastructure. Classification systems do not merely describe people and things. They organize work, authority, visibility, memory, and later action. A medical code can coordinate care while forcing a messy condition into a narrow label. A student-success category can help an adviser notice a student while turning a complicated person into a risk profile.

Rosa becomes legible through categories: enrolled, active, late, incomplete, aided, on track, at risk, full time, part time, first generation, online, in person. Some categories may help her. Some may harm her. Most are incomplete.

Classification also has a memory problem. Once a category enters a record, it can outlive the context that produced it. A missed week can become a pattern. A financial-aid problem can look like academic disengagement. A disability accommodation can become a hidden part of an institutional workflow. A note written by one adviser can shape how another adviser reads the student months later.

Accuracy matters, but a perfectly accurate record can still be ethically thin. Rosa did submit at 11:47 p.m. That is true. The question is what the system does with that truth.

This is also a dignity problem. A person has dignity not because every detail of their life can be recorded, but because they are more than any record. A category becomes morally dangerous when it becomes the only way a person can appear to an institution. The problem is not only that a label may be false. It is that a person may be forced to appear through a label too thin to carry their situation.

It can also be an epistemic-justice problem. Epistemic justice concerns whether people are treated fairly as knowers, interpreters, and explainers of their own lives. If Rosa’s data trail is treated as more authoritative than Rosa’s explanation, the institution may know something real and still misunderstand her. A system can make a person visible while making that person’s own account harder to hear.

When A Signal Becomes A Proxy

The central shift is from signal to proxy.

A signal is a trace the system treats as evidence. Rosa’s login time is a signal. A quiz attempt is a signal. A late submission is a signal. Location pings, source citations, purchases, messages, resume keywords, heart-rate readings, user ratings, prompts, and writing patterns can function the same way in other systems.

A proxy is a signal used as a stand-in for something harder to measure. Clicks may stand in for engagement. Course completion may stand in for readiness. Health-care spending may stand in for medical need. Zip code may stand in for risk. Prior salary may stand in for market value. Writing fluency may stand in for quality. Source count may stand in for research depth. User attention may stand in for satisfaction.

Proxies are often necessary. Institutions cannot directly measure every complex human concern. An instructor cannot see every minute of student study. A college cannot personally interview every student every day. A doctor needs tests, history, and records. Search engines need signals of relevance. Recommender systems personalize by treating some past behavior as evidence.

Proxy discipline begins when we ask how well the stand-in fits the thing we actually care about.

A proxy has to be judged in relation to the action it supports. A rough proxy may be acceptable for a low-stakes suggestion and unacceptable for a high-stakes gate. If a streaming service uses past viewing to recommend a movie, the cost of error is usually small. If a college uses a risk score to route advising, a rough proxy needs more care. If an employer uses a hiring model to screen candidates before a human sees them, the proxy carries heavier consequences. If a health system uses cost as a stand-in for need, the proxy can import unequal access to care into the decision.

Ethical burdens at different levels of proxy use
Proxy use Example Ethical burden
Low-stakes suggestion A recommender suggests a movie, playlist, or article Lower burden, though manipulation and filter bubbles can still matter
Supportive intervention An adviser receives an outreach flag Moderate burden: use the proxy to ask better questions, not to settle the judgment
High-stakes gate A model screens applicants, aid eligibility, health priority, or discipline High burden: stronger evidence, explanation, oversight, and appeal are needed
Field-shaping infrastructure A platform ranking, training dataset, or risk model becomes normal across a field Very high burden: the proxy may reshape what people make, learn, see, or become

Narayanan and Kapoor’s AI Snake Oil is useful for this point because they warn against treating prediction as understanding. A system may predict something useful from available signals while still failing to capture the thing people care about. Predicting who will click is different from knowing what helped someone learn. Predicting who is likely to miss class is different from knowing who needs support. Predicting which applicant resembles past hires is different from knowing who would flourish in the job.

Run the proxy question through Rosa’s week.

Suppose Rosa’s course has an early-alert dashboard. The dashboard combines login frequency, video views, due dates, quiz attempts, and late work. It produces a flag: “possible disengagement.” The flag could be helpful. It may prompt an adviser to send a friendly message before Rosa falls too far behind. It may help the instructor see that several students got stuck at the same point. Used that way, the proxy is a tool for attention.

The same flag can become morally weak very quickly. Rosa may have read offline, watched with captions downloaded earlier, or worked from a borrowed phone. Another student with the same activity pattern may be lost, exhausted, or ready to withdraw. The signal is identical. The human situation is different. A good system would treat the proxy as a reason to ask better questions. A bad system would treat it as a settled judgment.

Login frequency may be a useful signal. If Rosa has not logged in for two weeks, an adviser may have reason to check in. The signal could support care. A system that treats fewer logins as lower motivation may misread offline work, phone access, work schedules, or confidence. Late work may also stand in for confusion, fatigue, caregiving, illness, transportation, anxiety, overwork, weak planning, or a badly designed assignment sequence.

The same proxy can be useful in one context and reckless in another. A late submission may justify a friendly reminder. It should not automatically justify a judgment about character. A reading click may help an instructor identify a confusing page. It should not automatically become a measure of learning. A risk label may help a college offer support. It should not quietly become a lowered expectation.

This section sets up later chapters. Chapter 102 will ask what happens when proxies become opportunity gates. Chapter 315 will ask what happens when proxies steer delegated action. Chapter 319 will ask what happens when prior human work becomes training material for new outputs. Each question starts here, with the discipline of asking what the system is treating as evidence.

When Data Travels

Data becomes more ethically complicated when it moves.

A piece of information may make sense in one context and become troubling in another. A student tells an instructor about a family emergency. A patient searches for medication discounts. A person visits a place of worship. A teenager watches mental-health videos. A creator posts an image online. A programmer shares code in a public repository. A student asks a chatbot a personal question. The trace may have one meaning in the setting where it was produced and another meaning when it is aggregated, sold, scraped, used for training, or joined to other records.

Helen Nissenbaum’s theory of contextual integrity gives students a useful way to think about this. Privacy is not only secrecy. Social life depends on information flow. Teachers need some student information. Doctors need patient information. Employers need payroll information. Friends share stories. Families coordinate schedules. A violation can occur when information flows in a way that breaks the expectations and purposes of the original context.

“Public” is not a magic word. A post on a public forum may be visible, while its later use for model training, profiling, health inference, or ad targeting still requires judgment. A location ping may look harmless by itself, while a pattern of location pings can reveal sensitive routines. A student record may be legitimate for advising and illegitimate for unrelated vendor profiling. A health-app interaction may be useful for the service and troubling when it flows into advertising.

Regulators have been wrestling with these problems. The Federal Trade Commission’s 2024 report, A Look Behind the Screens, describes large-scale data practices among social media and video streaming services. FTC action involving Gravy Analytics and Venntel shows why location traces can become sensitive when they reveal visits to places such as medical facilities, religious sites, schools, or military locations. The FTC case timeline includes a January 2025 Final Consent Order and a press release announcing that the agency finalized an order prohibiting Gravy Analytics and Venntel from selling sensitive location data. The exact legal rules will continue to change, but the ethical pattern is stable enough for this course: data can change meaning when it travels.

Data travel is not always bad. Portability can be a good thing. A student should not have to repeat the same information to every office. A medical record can help clinicians avoid mistakes. A transcript can help a transfer institution understand prior learning. Content provenance can help audiences see where an image came from. A shared dashboard can help a team coordinate support.

Judge the movement by context, purpose, limits, and the people affected.

A Second Data Path: A Pharmacy Coupon App

Imagine someone searches for a discount on a medication through a coupon app. In the original context, the person may be trying to save money on health care. The data path may look helpful: a search becomes a coupon, the coupon lowers a price, and the person gets access to treatment.

Now trace what else may happen. The medication search can become a health-related trace. That trace may be linked to device identifiers, location, purchase behavior, or advertising categories. A company may infer a condition, a vulnerability, a likely future purchase, or a household pattern. The person may not expect the same information that helped them find a discount to travel into profiling, targeted advertising, brokerage, insurance inference, or employer wellness scoring.

The coupon example does not make coupon apps automatically wrong. It shows why context matters: a trace created for one practical purpose can become ethically different when it travels into another system with different incentives, different viewers, and different consequences.

From Records To Output Or Action

Records can sit quietly in a file. They can also become action.

Once data is stored and made comparable, systems can use it to classify, predict, rank, recommend, route, generate, and automate. A student’s data trail can support an advising alert. A platform profile can support recommendations or ads. A health record can support triage. A hiring system can rank applicants. A financial profile can influence offers. A large training set can help an AI model generate text, code, images, and analysis.

It helps to distinguish three downstream uses.

A prediction estimates something about a person, situation, or future event. It may estimate dropout risk, disease risk, likelihood of repayment, probability of clicking, likelihood of fraud, or chance that a source is relevant.

A generated output produces new text, image, audio, code, summary, label, recommendation, explanation, or classification from learned patterns. It may produce a draft email, an image, a research summary, a code suggestion, a chatbot answer, a tutoring hint, or a score explanation.

An action changes what happens next. It may send a reminder, route a patient, rank a resume, flag a transaction, change a price, withhold an offer, recommend a video, notify an adviser, generate a warning, or trigger a human review.

At this point, datafication becomes a technology of capacity. It can personalize feedback, adapt services, detect anomalies, reveal patterns, audit institutional processes, and coordinate action across people and tools. A college may use records to find students who need help earlier than an instructor could alone. A workplace may use logs to improve safety. A scientist may use data to discover a pattern no human observer could find unaided. A public agency may use records to allocate resources more accurately.

The same path can become a gate. A gate is any decision point that shapes access to attention, help, credit, opportunity, care, credibility, visibility, or second chances. A risk score can move a student into support or stigma. A ranking can move a resume to the top or bottom of a list. A price model can alter what someone is offered. A recommendation system can decide what a creator’s audience sees. A generated summary can shape what a reader believes about a source.

This is why the data path matters before the ethical verdict. If the records are thin, the categories are crude, the proxy is weak, and the action is high-stakes, the system deserves more scrutiny. If the records are limited, the purpose is clear, the proxy is modest, the action is supportive, and the affected person can challenge the result, the same kind of datafication may be easier to defend.

The chapters that follow will slow down at different points along this path. Chapter 102 asks how data-driven gates become biased or unjust. Chapter 315 asks what happens when systems delegate action to automated loops, agents, robots, and dashboards. Chapter 319 asks how cultural and creative work becomes training material for generated outputs.

Chapter 91 gives the first move. Do not start with the final output. Start with the path.

Cultural Work Can Become Data

Datafication is often described as a privacy issue because many examples involve personal information: location, browsing, purchases, grades, health, messages, and biometrics. That is only part of the story. Datafication also reaches cultural and creative work.

Books, images, songs, code, scientific records, forum posts, captions, product reviews, videos, prompts, and ordinary web pages can become training material. This is one reason generative AI systems can translate, summarize, code, answer questions, generate images, and assist research. They are built from large collections of prior human expression and structured feedback.

That capability is real. A model that can translate quickly can widen access. A model that can summarize a technical report can help a beginner enter a field. A model that can generate code examples can help a student practice. A model trained on scientific data can help researchers notice patterns. A model that can produce alt text, captions, drafts, and examples can make creative and academic work more accessible.

The ethical questions remain. Who created the material? Was it licensed, scraped, purchased, contributed, or generated synthetically? Was the material public in a way that makes this use appropriate? Did creators, users, or communities have any way to refuse? Who cleaned, labeled, moderated, or rated the data? What kinds of work were excluded? What kinds of work were overrepresented? What field-level effects follow when generated outputs compete with or reshape the work that trained the system?

The U.S. Copyright Office’s 2025 AI initiative page says that Part 3 of its copyright-and-AI report, Generative AI Training, remains a pre-publication version, with final publication expected in the future and no substantive changes expected in the analysis or conclusions. This chapter does not need to decide the legal question. Students should separate several questions that often get blurred together: whether a work was publicly available, whether it was legally usable, whether its use was ethically appropriate, whether the model output is copyrightable, whether creators were harmed, and whether the field gained something valuable.

Provenance helps, but it does not settle the moral question. Provenance means the record of where something came from and how it changed. Standards such as C2PA Content Credentials try to preserve information about the origin and history of digital media. The National Institute of Standards and Technology’s report on synthetic content transparency treats provenance, watermarking, metadata, and labeling as partial tools for understanding origin and history, while also emphasizing that these tools vary in robustness and require people and institutions to use them well. Those tools can help. They do not prove that a use was fair, accurate, respectful, legal, or trustworthy.

Rosa’s Canvas record and a writer’s public essay look very different. They still share one question: what path did the data take before it became useful to a system?

What Makes Datafication Useful And More Defensible?

At this point, a fair objection should be on the table. If data can help colleges support students, doctors catch problems, researchers discover patterns, public agencies allocate resources, platforms improve access, and creators build new tools, why treat datafication as ethically serious?

The answer is that usefulness raises the stakes. A system that does nothing useful is easy to reject. The harder systems are useful enough to adopt and powerful enough to harm.

Several ethical stakes now come into view.

The first stake is knowledge. A data system offers a way of knowing a situation, but every way of knowing has limits. Rosa’s data trail may reveal a pattern her instructor would miss. It may also hide the reason for the pattern. That is an epistemic problem, meaning a problem about what counts as knowledge. A system can know something real and still know it too thinly for the decision being made. An epistemic injustice occurs when a system makes it harder for someone to be understood, believed, or interpreted fairly.

The second stake is power. Datafication changes who can see, classify, compare, and act. A student may not know which traces are being collected or how a dashboard reads them. A platform may know more about a user’s habits than the user can see in return. A company may turn millions of creative works into training material while individual creators struggle to discover whether their work was included. The ethical question is not only whether the record is accurate. It is also who gains practical power from the record.

The third stake is agency. People need room to explain, correct, refuse, or redirect how records are used. If Rosa receives a support message because the system noticed a risk pattern, the data trail may serve her agency. If the same pattern quietly lowers expectations or sends her into an opaque category she cannot challenge, it works against her agency. Contestability matters because human beings are not only data subjects. They are people trying to act within systems that increasingly act on them.

The fourth stake is justice. Data-driven systems distribute attention, opportunity, cost, suspicion, and help. A weak proxy can make the wrong students visible, leave others unsupported, or turn a past pattern into a future barrier. A useful data system can also reveal an inequity that an institution was ignoring. Justice asks who gets helped by the data path, who absorbs the errors, who is missing from the data, and who has enough standing to question the result.

The fifth stake is care. Many data systems are adopted because institutions want to help people earlier or more consistently. That can be a real good. A warning sign can help a nurse notice deterioration, an adviser notice financial-aid trouble, or a social-service office notice an unmet need. But care depends on attention to context. A care system becomes colder when it replaces listening with labels, treats a proxy as a diagnosis, or makes the most vulnerable people responsible for correcting records they cannot even see.

Datafication is more defensible when the purpose is clear. If a college collects activity data to help instructors identify confusing parts of a course, that purpose is easier to evaluate than a vague claim about “improving student success” with no limits. Datafication is more defensible when the data collected is proportionate to that purpose. A reminder system may need deadlines and submission status. It probably does not need every unrelated student trace a vendor can collect.

It is more defensible when the data path respects context. Advising data should not silently become advertising data. Health-related searches should not casually become marketing profiles. A public forum post should not be treated as if every future use is equally expected. Context allows information to move while keeping the original relationship and purpose ethically visible.

It is more defensible when the proxy is disciplined. A login can be a signal for outreach. It should not become a full judgment about motivation. A grade can be a signal for mastery. It should not become the whole story of a student’s ability. A model’s fluency can be a signal that the answer is readable. It should not become proof that the answer is true.

It is more defensible when the affected person has some route to correction or appeal. Contestability matters because data systems make mistakes and because accurate records can still be misleading. Rosa should have some way to explain why the late work happened, correct a wrong record, ask what a label means, or challenge a decision that affects her path.

It is more defensible when the benefits and burdens are visible. Catherine D’Ignazio and Lauren Klein’s Data Feminism gives students a practical question set for this: who has power, who is missing, whose labor is hidden, who benefits, and who can intervene? Linnet Taylor’s article on data justice pushes in the same direction by connecting datafication to how people are made visible, represented, and treated.

These questions do not require students to reject data-driven systems. They require students to judge them at the right level. The issue is rarely one data point by itself. The issue is the path from trace to record to category to proxy to output or action.

A Short Practice: Trace One Data Path

Choose one data-driven system you use or might encounter in your field. It could be a learning platform, a fitness app, a hiring screen, a recommendation system, a health portal, a navigation app, a chatbot, a plagiarism detector, a scheduling system, a customer-service tool, or a generative AI model.

Write a short data-path note. Keep it concrete:

  1. What human activity, work, text, image, movement, body, place, or event became data?
  2. What trace was captured?
  3. What was omitted or simplified?
  4. What record, profile, dataset, or log stored the trace?
  5. Who made the categories?
  6. What useful capability did the data make possible?
  7. What signal became a proxy?
  8. What prediction, ranking, recommendation, generated output, or action followed?
  9. Who benefited?
  10. Who was affected or missing?
  11. Who could inspect, explain, correct, appeal, or refuse?

If you use Rosa’s case, the note might begin like this:

The system turns logins, clicks, video views, quiz attempts, submissions, grades, advising notes, and aid status into a student data trail. It may use timestamps and activity patterns as proxies for engagement or risk. That can support early outreach, but it may miss offline reading, work schedules, caregiving, and financial-aid confusion. The data path is more defensible if the result is supportive, limited, explainable, and correctable.

Here is a non-school example:

A fitness app turns steps, heart rate, sleep timing, workout logs, and location into a health data trail. It may use movement patterns as a proxy for wellness, effort, or risk. That can support reminders, coaching, personal insight, or medical conversation. It becomes less defensible if the data travels into advertising, insurance inference, employer wellness scoring, or opaque profiling without correction, refusal, or clear limits.

This kind of tracing does not settle the case. It gives students a better starting point for judgment.

The goal is disciplined trust. A data system can help people see what they would otherwise miss. It can also make the wrong thing easier to act on. Ethical data systems should make people, institutions, and decisions visible to one another in ways that can be questioned.

What To Keep

Trace before verdict. Before deciding that an AI system is fair, biased, helpful, or dangerous, ask how the data path was built.

Datafication is useful because it makes patterns visible and lets systems coordinate action at scale. That usefulness is exactly why it needs ethical judgment.

Data is made. It is shaped by recording choices, categories, formats, labor, storage, access, and purpose.

A record is not the whole person, even when the record is accurate. It preserves some traces and drops others.

A category makes someone or something legible. Legibility can support care and coordination. It can also flatten context and strengthen control.

A proxy is a claim about fit. It may be strong enough for a reminder and too weak for a high-stakes decision.

Data changes meaning when it travels. Context, purpose, and limits matter.

Public availability is not permission. Cultural work can become training data, but legality, ethics, consent, labor, and context are different questions.

A defensible data path has a clear purpose, a proportionate record, a context-respecting flow, a disciplined proxy, visible benefits and burdens, and some route for correction or appeal.

References

Scholarly layer

Person record