Do AI Agents Replace Freelancers? What 4 Years of Evidence Say (ChatGPT, Claude, Upwork, Jobbit)
What four years of evidence say about AI agents and freelance work: which tasks were substituted, where people still win, and how a business should split work between agents and humans.

Every business owner is now asking a version of the same question: which work can be handed to an AI agent, and which still needs a person? For once, the question has a laboratory. Since late 2022, the online freelance platforms that host millions of writers, translators, designers and developers have also hosted a natural experiment in what happens when a capable substitute for routine knowledge work arrives overnight, at near-zero marginal cost, into a market where every job is posted, priced and timestamped.
This case study treats that experiment the way a scholar would treat any other body of evidence: what exactly was measured, how well, what the results show, and what they cannot yet show. It draws on peer-reviewed studies and working papers from researchers at Stanford, MIT, Harvard Business School, the University of Chicago and UCLA, on randomised field experiments inside BCG and Alibaba, on capability evaluations from METR and OpenAI, and on the platforms' own data from Upwork and Fiverr. The period covered runs from November 2022, when ChatGPT was released, to August 2026, when the most recent of these studies was updated.
The short answer to the title is that agents have so far replaced tasks rather than freelancers, and that the replacement has been narrow, fast and uneven. The long answer is more useful, because it tells a business where the boundary currently sits and how to work on both sides of it.
Summary of findings
Five conclusions survive a sceptical reading of the evidence.
First, substitution is real and concentrated. Within eight months of ChatGPT's release, postings for automation-prone writing and coding work on a leading global platform fell by about 21% relative to manual-intensive work, and image-generation tools cut postings for image work by about 17%. On Upwork, credentials and reputation have become measurably less predictive of who gets hired in AI-exposed categories, while price has become more predictive.
Second, the effect on the wider labour market remains small. The most precise study, on Danish administrative records, rules out effects on earnings or hours larger than 2% two years after ChatGPT. The exception is sharp: in the United States, employment of workers aged 22 to 25 in the most AI-exposed occupations now sits about 19% below where it would be had it kept pace with less exposed peers.
Third, the productivity gains from working with AI are large but uneven. They are biggest for less experienced workers on routine tasks and can turn negative for experts on tasks outside the tool's competence. In one rigorous trial, experienced developers were 19% slower with AI tools while believing they were 20% faster.
Fourth, agents have crossed from assistants to workers. The length of task a frontier system can complete has been doubling every few months, and blind expert graders now judge the best models' deliverables as equal to or better than a professional's in nearly half of cases. But once review and correction time is counted, the measured saving collapses from roughly ninety-fold to between 12% and 39%, and even the best workplace agents still occasionally make small mistakes with irreversible consequences.
Fifth, the rational response, and the one the platforms are converging on, is a division of labour: agents for routine, reviewable digital work; people for judgment, accountability, presence and anything irreversible; and a deliberate hand-off between the two.
The case: why freelance platforms are the right laboratory
Labour economists usually work with annual surveys and broad occupational codes. Online freelance platforms offer something rarer: work broken into individual tasks, each with a posting date, a price, a category, a buyer, a seller and an outcome. When a technology arrives that can perform some of those tasks, the change shows up within weeks in the flow of postings, in who wins the contract and at what rate. The market also adopts new tools instantly, because there is no procurement cycle between a freelancer and a browser tab.
The same features make the laboratory imperfect. Platforms record posted demand, not all demand; a client who now does the work in-house with an agent disappears from the data rather than appearing as a substitution. Categories shift, so a decline in "writing" may partly be a relabelling into "AI content editing". Most of the large datasets are American, and platform freelancers are not representative of the workforce. These limits matter for interpretation, and they are returned to below, but they do not make the evidence useless. They make it a leading indicator rather than a census.
The stakes are not small. IPSE estimates the UK freelance workforce at about 2.05 million people, roughly half of the country's 4.2 million solo self-employed, using a definition confined to the three highest-skilled occupational groups. Those are precisely the occupations that AI applicability studies rank highest.
A timeline of the evidence, 2022 to 2026
| Year | Study | Setting and design | Headline finding |
|---|---|---|---|
| 2023 to 2024 | Hui, Reshef and Zhou, Organization Science | Large freelance platform, difference-in-differences around ChatGPT and image models | Exposed freelancers lost both jobs and earnings; strong past performance did not protect them |
| 2023 | Dell'Acqua and colleagues, Harvard Business School with BCG | Randomised trial, 758 consultants, GPT-4 | Inside the tool's range: 12.2% more tasks, 25.1% faster, 40% higher quality. Outside it: 19 points less likely to be correct |
| 2024 to 2025 | Demirci, Hannane and Zhu, Management Science | Leading global platform, job postings | Automation-prone writing and coding posts down 21% within eight months; image posts down 17% |
| 2025 | Brynjolfsson, Li and Raymond, Quarterly Journal of Economics | 5,179 customer support agents, staggered rollout | Productivity up 14% on average, 34% for novices, little for experts |
| 2025 | Humlum and Vestergaard, NBER, revised 2026 | Denmark, adoption surveys linked to administrative records | Earnings and hours effects within 2% two years on; work reorganised instead |
| 2025 | METR | Randomised trial, 16 experienced developers, 246 real tasks | 19% slower with AI tools; participants believed they were 20% faster |
| 2025 | OpenAI, GDPval | 44 occupations, blind expert grading | Best model's deliverables win or tie 47.6% of the time; savings fall to 12% to 39% once review time is counted |
| 2025 to 2026 | Brynjolfsson, Chandar and Chen, Stanford | ADP payroll records to June 2026 | Workers aged 22 to 25 in exposed occupations 19% behind peers; no economy-wide displacement |
| 2026 | Siddiq and Zhang, UCLA | 2.26 million Upwork contracts, 2021 to 2026 | Credentials and reputation 7.8% less predictive of hiring in exposed categories; price more predictive |
| 2026 | Wang and colleagues, Alibaba | Randomised field experiment, agentic AI with human supervisors | Shorter chats but lower ratings when the agent handled them; humans rescued technical failures better than emotional ones |
| 2026 | Styles and Miller, WorkBench revisited | Workplace agent benchmark, 2024 versus 2026 | Task completion from 43% to 98%; harmful actions from 26% to 1.9%, not zero |
| 2026 | Upwork, Future Workforce Index | Survey of 2,400 US skilled workers plus platform data | AI work pays 34% more per hour; simple generative work pays 13% less per contract than a year earlier |
Finding 1: substitution is real, fast and narrow
The earliest rigorous evidence came from Xiang Hui, Oren Reshef and Luofeng Zhou, whose study of a large online labour platform was published in Organization Science in 2024 and later won Washington University's Olin Award for research with practical impact. Comparing freelancers in occupations exposed to ChatGPT, DALL-E 2 and Midjourney with those in less exposed occupations, they found declines in both the number of jobs and monthly earnings for the exposed group. The finding that unsettled the field was the one about quality. Freelancers with strong past ratings and long track records were not shielded; if anything, the authors report suggestive evidence that top performers were hit harder. The most plausible reading is that in routine text and image work, an excellent human output and an adequate machine output had become close substitutes in the eyes of buyers, so the premium for excellence eroded first.
Ozge Demirci, Jonas Hannane and Xinrong Zhu then measured the demand side directly. Using postings from a leading global freelancing platform, in a paper published in Management Science in 2025, they found a 21% fall in postings for automation-prone writing and coding jobs relative to manual-intensive jobs within eight months of ChatGPT's release, and a 17% fall in image-creation postings after the arrival of image-generating models. Two secondary results matter as much as the headline. The postings that remained in the exposed categories were more complex and paid more, and competition for them intensified. Substitution, in other words, did not empty the category; it removed the bottom of it.
Why writing, translation and image work first? Microsoft Research's analysis of 200,000 anonymised Bing Copilot conversations, published in 2025, gives the mechanism a shape. The occupations with the highest AI applicability were those whose work consists of gathering, producing and communicating information; interpreters and translators topped the list, with 98% of their work activities overlapping with tasks the assistant already performed well. Applicability is not displacement, and the authors are careful to say so, but the ranking predicts almost exactly which freelance categories the platform studies found to be shrinking.
The most recent and, for a business reader, the most instructive evidence comes from Auyon Siddiq and Niuniu Zhang at UCLA Anderson. Their 2026 working paper, "Human Capital, AI, and Labor Commoditization", follows nearly 50,000 Upwork freelancers who were active before ChatGPT across 21 quarters from early 2021 to early 2026, covering 2.26 million completed contracts. Rather than counting jobs, they ask what predicts who gets hired. In more AI-exposed categories such as translation and writing, the combined predictive weight of a freelancer's self-presentation, credentials and reputation fell by about 7.8% relative to less exposed categories such as photography, and the weight of price rose. By the final year the human capital measures had lost about 10.1% of their importance and price had gained about 1.8%. The advantage held by experienced freelancers narrowed by 6.2% and then by 10.3%, and in the most exposed work, cheaper freelancers won 7.9% more contracts. Overall demand in exposed fields fell by about 7%. The authors' word for this is commoditisation: when a buyer can get a passable first draft from a model, what they will pay a human for shifts from "who are you" to "how much".
Finding 2: the aggregate effect is still small, with one loud exception
If the platform evidence were the whole story, national statistics should by now show it. Mostly, they do not, and the studies that fail to find large effects are among the best designed.
Anders Humlum and Emilie Vestergaard linked large-scale adoption surveys to Danish administrative labour records, covering 25,000 workers and 7,000 workplaces in eleven exposed occupations. Their NBER working paper, first circulated in 2025 and revised in March 2026, documents rapid adoption: most employers in exposed occupations had launched chatbot initiatives, workers reported productivity benefits, and new AI-related tasks were widespread. Yet the difference-in-differences estimates on earnings and recorded hours are precise nulls, ruling out effects larger than 2% two years after ChatGPT, at both the worker and the workplace level. What moved was the structure of work: employers absorbed the technology through task reorganisation, including new tasks in content generation, AI oversight and AI integration, and adopters drifted into higher-paying occupations where the tools are more useful, though in numbers too small to shift averages. The earlier version of the paper put self-reported time savings at under 3% of work hours, a fraction of the gains, often above 15%, documented by controlled trials in the same occupations. The gap between what AI can do in a trial and what it changes in a payroll is the central puzzle of this literature, and Denmark is where it is measured most cleanly.
The Budget Lab at Yale reaches a similar conclusion for the United States by a different route. Tracking the occupational mix of employment and applying synthetic difference-in-differences to AI-exposed occupations, its 2025 and 2026 updates find no clear AI footprint. Early 2026 shows low layoffs but also low hiring, and the authors argue that a technology with widespread effects would be expected to take longer than the roughly three years since ChatGPT to show them.
The exception is documented by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen at Stanford's Digital Economy Lab, whose paper "Canaries in the Coal Mine?" uses ADP payroll records covering millions of American workers. In its August 2026 update, with data through June 2026, employment of workers aged 22 to 25 in highly AI-exposed occupations stood about 19% below where it would be had it tracked similarly aged workers in less exposed occupations, up from a 15% gap a year earlier. In levels, employment of that age group in the two most exposed quintiles fell about 11% between November 2022 and June 2026 while employment in the three least exposed quintiles grew about 10%. Experienced workers show no comparable gap, and the authors state plainly that they do not see widespread, economy-wide displacement. The declines concentrate in occupations where AI use automates work rather than complementing it.
These results are consistent with each other once one stops expecting AI to behave like a recession. Substitution at the task level appears first where contracts are shortest and re-priced most often: freelance gigs and entry-level hiring. The freelance market is the canary because it re-prices every week; the payroll is the mine, and the air there has changed least so far.
Finding 3: augmentation gains are large but uneven, and the frontier is jagged
The complement to substitution is augmentation, and here the experimental evidence is unusually strong because randomisation is possible.
Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered introduction of a generative AI assistant to 5,179 customer support agents, in a paper published in the Quarterly Journal of Economics in 2025. Issues resolved per hour rose 14% on average, with a 34% gain for novice and lower-skilled agents and little change for experienced, high-skilled ones. The tool appeared to encode the practices of the best agents and pass them to the newest, moving them down the experience curve faster.
Fabrizio Dell'Acqua and colleagues at Harvard Business School, working with BCG, randomised 758 management consultants across realistic tasks with and without GPT-4, in a study now published in Organization Science. On eighteen tasks inside the model's competence, consultants with AI completed 12.2% more tasks, 25.1% faster, with quality rated 40% higher by blind evaluators. On a task deliberately chosen to sit outside that competence, consultants with AI were 19 percentage points less likely to reach the correct answer. The authors' phrase for this boundary, the jagged technological frontier, has become the standard description of the problem: the tool's competence does not run along lines a user can see, and users are poor at telling which side of the line they are on.
A field experiment inside Alibaba's customer service operations, by Xiao Ni, Tianjun Feng, Lauren Xiaoyuan Lu and co-authors, replicated the distributional pattern at scale: 5,940 new agents, roughly 2.56 million chats, 390,000 customer ratings. Low performers improved most in both speed and quality; top performers gained little speed and lost quality on both subjective and objective measures. Assistance that raises the floor can lower the ceiling.
Then there is METR's randomised trial of experienced open-source developers, published in July 2025, which is the most cited cautionary result of the period and deserves a precise statement. Sixteen developers worked 246 real tasks in repositories they maintained, averaging over 22,000 stars and a million lines of code, with tasks randomly allowed or denied AI tools, mainly Cursor with Claude 3.5 and 3.7 Sonnet. With AI they took 19% longer. Before the study they had forecast a 24% speed-up; afterwards they estimated they had been 20% faster. The perception gap is as important as the productivity result. It suggests that self-reported gains, which underlie most industry surveys, are biased upwards precisely where expertise is highest. The implications for software work specifically are discussed in will AI replace software developers?
A 2025 review by R. Maria del Rio-Chanona, Ekkehard Ernst, Rossana Merola, Daniel Samaan and Ole Teutloff pulls these strands together across the literature: productivity gains of roughly 20% to 60% in controlled trials and 15% to 30% in field experiments, larger for novices on simple tasks, mixed on complex tasks, with digital trace data showing substitution in writing and translation alongside rising demand for AI skills and mild but growing evidence of declining demand for novices.
Finding 4: agents crossed from assistant to worker, and the reliability gap is the whole story
Everything above concerns AI as an assistant to a human who remains in the loop for every step. The change since 2025 is that systems now execute multi-step work with the human checking at the end, or not at all, and that changes which evidence matters.
METR's time-horizon measure asks how long a task, expressed in the time a human professional would need, a system can complete with a 50% success rate. In the January 2026 revision of the methodology, with a suite expanded to 228 tasks including 31 of eight hours or longer, the long-run doubling time was estimated at about 196 days, with faster doubling of 131 days since 2023 and 89 days since 2024. The best measured model at that point, Claude Opus 4.5, had a 50% horizon of about 320 minutes, with a wide confidence interval of roughly three to twelve hours. The authors are explicit that a 50% success threshold is generous for real work, that the trend is sensitive to task composition, and that most long tasks use estimated rather than measured human times. Taken with those caveats, the direction is unambiguous: the length of digital work that can be delegated has been growing exponentially.
OpenAI's GDPval evaluation, released in late 2025, tests whether that delegated work is any good. Tasks were written by professionals in 44 occupations across nine sectors, averaging 14 years of experience, and graded by blinded expert pairwise comparison against the human deliverable. The best model, Claude Opus 4.1, produced deliverables graded better than or equal to the expert's in 47.6% of cases. The number that a business should remember is a different one. The naive comparison of machine time to human time yields ratios near ninety-fold, but once the paper accounts for a human reviewing and, where needed, redoing the work, the estimated saving falls to between 1.12 and 1.39 times depending on strategy. Review is where the cost went. That is the same mechanism Dell'Acqua found in consultants and METR found in developers, now measured at the level of the deliverable.
Workplace benchmarks show both the speed of progress and the residue. TheAgentCompany, from Carnegie Mellon, dropped agents into a simulated software firm with real instances of GitLab, OwnCloud, Plane and RocketChat and 175 tasks; the best agent completed about a quarter of them at launch in late 2024 and about 30% during 2025. Olly Styles and Sam Miller revisited their WorkBench benchmark in June 2026 and reported that the best agent's completion rate had risen from 43% for GPT-4 in March 2024 to 98% for Claude Fable 5, while the share of runs with unintended harmful actions fell from 26% to 1.9%. They also report that frontier models still make basic mistakes that occasionally cause irreversible harm. Two per cent is very good; it is also roughly one run in fifty doing something it should not.
The 2026 Alibaba field experiment, by Yiwei Wang, Chuan Zhu, Tianjun Feng, Lauren Xiaoyuan Lu and Bingxin Jia, is the first randomised evidence on agentic AI with humans in the loop in a live service operation. Treated workers supervised an agent that handled eligible customer chats while they worked the rest by hand. Average chat duration fell with limited effect on customers coming back with the same problem, but chats the agent handled received substantially lower ratings. Human intervention recovered the situation far better when the failure was technical, an unresolved issue, than when it was emotional, a frustrated customer, and it worked best when it came early. Supervising workers also paid more attention to the chats they still handled themselves, a positive spillover. This is the most direct evidence yet on how the hand-off between agent and person should be designed: watch early, intervene on facts, and do not leave feelings to the machine.
Adoption data confirm the shift is not confined to laboratories. Anthropic's Economic Index for March 2026 reports that 49% of occupations have seen at least a quarter of their tasks performed with Claude, and that coding work has migrated from conversational use towards automated API workflows, with the coding share rising 14% in the API and falling 18% in the chat product since August 2025. Lightcast's contribution to the 2026 Stanford AI Index finds AI skills in 2.5% of all US job postings, and postings mentioning an "agentic AI" skill cluster rising from 0.06% in 2024 to 0.23% in 2025, roughly 90,000 postings. And UpBench, a 2025 benchmark built from real, completed jobs on the Upwork marketplace, has expert freelancers decompose each job into verifiable acceptance criteria against which agents are scored. Grading agents with the rubric a client would apply to a freelancer is a signal of where the supply side is expected to go.
Mechanisms: why the pattern looks like this
Four mechanisms account for most of what the evidence shows, and they are worth stating because they predict what happens next.
The first is commoditisation of routine competence. When a model produces an acceptable draft of a product description, a translation or a boilerplate module, the buyer's willingness to pay for a human's credentials in that category falls, and price becomes the deciding variable. Siddiq and Zhang measure exactly this shift, and Demirci and colleagues show its mirror image: the work that stays human becomes more complex and better paid, and more contested.
The second is verification cost as the binding constraint. Production has become cheap; checking has not. GDPval's collapse from ninety-fold to under 1.4-fold once review is counted, METR's slowed experts, and the jagged frontier all describe the same economy. Whoever can verify an output cheaply captures the gain; whoever cannot pays for it in rework or in undetected error. This is why gains flow to novices on tasks with easy checks and stall for experts on tasks where checking costs as much as doing.
The third is the apprenticeship gap. The tasks an agent does well are the tasks firms used to give to junior staff and junior freelancers so they could learn. Brynjolfsson, Li and Raymond show that assistance helps novices most; Brynjolfsson, Chandar and Chen show that firms nonetheless hire fewer of them. Both can be true: a firm that needs fewer juniors to produce the same output hires fewer juniors, even if each is more productive. The long-run cost, a thinner pipeline of experienced people, is not yet visible in any dataset and is the most important unmeasured variable in this field.
The fourth is the persistence of complementary human tasks. Demirci's comparison group of manual-intensive work did not fall. The Alibaba experiment shows emotional failures resisting machine recovery. Accountability, presence, physical work and relationships remain human for reasons that have nothing to do with model capability and everything to do with what a customer will accept.
A sceptical reader should hold two alternative explanations alongside these. The entry-level decline in the United States coincides with higher interest rates and the unwinding of pandemic-era hiring in technology, which independently hit the same cohort; the Stanford authors address this with within-firm comparisons and by contrasting automating with augmenting occupations, but the confound cannot be fully excluded. And every "exposure" measure is itself model-derived, so studies partly test the assumptions of the classifier. Denmark's collective bargaining institutions may also damp wage responses that would appear faster in a looser labour market.
How the platforms responded
The marketplaces read the same signals and moved. In April 2025, Fiverr's chief executive Micha Kaufman wrote to staff that "AI is coming for your jobs", including his own, and urged them to master tools such as Cursor and Legora; in September 2025 the company cut about 250 roles, close to 30% of its workforce, to become what it called an AI-first company, having launched its Fiverr Go creator tools that February. Upwork has been turning its assistant, Uma, into what it describes as an always-on work agent that takes actions on behalf of clients and freelancers, and its 2026 Future Workforce Index reframes the successful freelancer as an "AI orchestrator" who connects tools to domain expertise and applies judgment. The same report gives the price signal in one line: freelancers doing AI work earned 34% more per hour than those who did not, complex AI-augmented work saw earnings rise 45% year on year, while generative AI and creative production work grew 90% in contract starts as per-contract earnings fell 13%. Volume moved to the simple end; money moved to the complex end.
Jobbit, the platform on which this case study is published, was designed around the division of labour the evidence points to rather than around either side of it. The agent takes the reviewable digital work: building and deploying a website or web app, drafting documents and presentations, running research, setting up automations, generating images and video. The user can watch it work in the agent browser and take control at any moment, which is the early-intervention pattern the Alibaba experiment found most effective. For work an agent does badly or should not do alone, physical and local jobs, accountable expert review, anything a customer needs a person for, the platform routes to a moderated human network on pro.jobbit.uk. Jobbit has not published trial data of its own, and this study makes no claim about measured outcomes on it; the claim is one of fit between a design and the evidence. Readers who want the mechanics can start with what agentic AI actually is, the guide to watching and taking control of the agent browser, and the practical list of AI use cases for a small business.
What a business should do with this evidence
The findings translate into a delegation rule that a business owner can apply without reading a single paper.
Give an agent the work that is routine and reviewable: first drafts, summaries, research briefs, standard documents, a website or internal tool you can click through and check in an afternoon. This is where every study finds the largest gains, especially where the business currently does the task slowly or not at all, because the novice effect applies to firms as much as to people. A practical starting sequence is set out in how to automate your business with AI agents.
Keep a person on the work that is irreversible, emotional, physical or accountable: anything that sends money, deletes data, touches a customer who is upset, or requires someone to stand behind the result. The Alibaba experiment and the WorkBench residue point the same way.
Price in review time before deciding an agent is cheaper. GDPval's honest arithmetic, a saving of 12% to 39% rather than ninety-fold, is the number to plan with. If nobody in the business can check the output, the saving is not real; the answer in that case is to have the agent do the work and a human expert review it, which is the arrangement the freelance data show buyers already paying a premium for.
Expect the frontier to be jagged and to move. A task the agent could not do in January may be routine by September, and the reverse is not true. Test, keep a record of what was checked and what failed, and revise the split every quarter.
The working rule from four years of evidence: delegate what is routine and checkable, keep a person on what is irreversible or emotional, and treat review time as part of the price.
Limitations and open questions
Most of the large datasets are American, and platform data measure posted demand rather than all work performed. Exposure measures are model-derived, and results partly inherit the classifier's assumptions. Horizons are short: two years in Denmark, under four in the United States, and the Yale authors' point that widespread effects may take longer is well taken. Benchmarks are not jobs; GDPval tasks are self-contained deliverables and WorkBench is a synthetic office, whereas real work is interrupted, ambiguous and political. Capability is moving quickly enough that any figure in this study is a snapshot, and the studies most likely to be superseded are the capability numbers, not the labour ones. Finally, the effect that matters most for the next decade, whether a thinner apprenticeship pipeline degrades the supply of experienced professionals, cannot yet be measured at all.
Conclusion
Do AI agents replace freelancers? So far they replace tasks, and they do it where a task's output can be produced and judged as routine. That has been enough to shrink the market for commodity writing, translation and image work, to weaken the hiring premium for credentials in those categories, and to cut entry-level hiring in exposed occupations sharply, while leaving earnings and hours across whole economies essentially unchanged. The freelancer who prospers in the data is the one selling judgment, accountability and the ability to direct and check machine work; the business that benefits is the one that learns to delegate the routine and to review it, and that keeps a person on everything that cannot be undone. The market is not losing one side of the transaction. It is reorganising around the hand-off between them, and the platforms that make that hand-off easy are the ones the evidence favours.
Frequently asked questions
Do AI agents replace freelancers?
Not as a category. The best evidence from 2022 to 2026 shows AI replacing specific tasks, chiefly routine writing, translation, image creation and simple coding, where job postings fell by roughly a fifth relative to manual work within months of ChatGPT. Demand for complex work in the same categories rose in pay and in competition, and national data show no economy-wide displacement. The clearest losers are entry-level workers in exposed occupations, not experienced freelancers.
Which freelance jobs are most affected by AI?
Writing, translation, content editing, image and design production, and routine coding. Studies of Upwork and other platforms find these categories lost postings and saw credentials matter less and price matter more when clients hire. Work that is physical, local, relationship-based, or that carries accountability for the result has been largely unaffected so far.
Is it cheaper to use an AI agent than to hire a freelancer?
For routine, checkable work, usually yes, but by less than headline comparisons suggest. OpenAI's GDPval evaluation found models produce professional deliverables roughly ninety times faster in raw terms, yet once human review and correction are counted, the saving falls to about 12% to 39%. If nobody in the business can review the output, the realistic option is an agent producing the work and a human expert checking it.
What work should a small business still give to a person?
Anything irreversible, emotional, physical or accountable: payments and deletions, upset customers, on-site jobs, and results someone must stand behind. A 2026 field experiment at Alibaba found agent-handled service chats were faster but rated lower, and that human intervention rescued technical failures far better than emotional ones. Platforms such as Jobbit pair an agent for digital tasks with a human network for exactly this remainder.
Why do studies disagree about AI and jobs?
They mostly measure different things. Controlled experiments measure what AI can do on selected tasks and find large gains; platform data measure demand for specific task categories and find sharp local declines; payroll and administrative data measure economy-wide employment and earnings and find little change so far. The picture is consistent once task-level substitution, uneven augmentation and slow aggregate adjustment are read together.