Business & Generative AI Conference
Keynote Speakers

Angela Duckworth
Rosa Lee and Egbert Chang Professor, University of Pennsylvania; Faculty Co-Director, Behavior Change for Good Initiative; New York Times Bestselling AuthorMore
Angela Duckworth is the Rosa Lee and Egbert Chang Professor of Psychology at the University of Pennsylvania and faculty co-director of the Behavior Change for Good Initiative. Angela is the #1 New York Times bestselling author of Grit: The Power of Passion and Perseverance. Her new book, Situated: Find the People and Places that Bring Out Your Best, comes out from Scribner on September 1, 2026.

Faye Iosotaluno
President & COO, Entangled Publishing; former CEO, TinderMore
Faye is a seasoned technology and consumer executive known for transforming global businesses, building category-defining brands, and driving growth at the intersection of technology, culture, and consumer behavior. She is currently President and Chief Operating Officer of Entangled Publishing, where she is helping lead the company’s next stage of growth by building on its successful publishing foundation, including the global bestselling Fourth Wing franchise, and expanding into new direct-to-consumer products, platforms, and experiences that connect stories, authors, and audiences in new ways. Entangled has been recognized by TIME as one of the TIME100 Most Influential Companies and by Fast Company as one of the World’s Most Innovative Companies.
Previously, Faye served first as COO and then as CEO of Tinder, where she led a broad brand and product transformation, repositioned the platform to resonate more deeply with women, advanced product innovation, and oversaw major business functions including member experience, strategy, finance, and international operations. She joined Tinder after four years at Match Group as Chief Strategy Officer, leading corporate strategy and new business growth across its global portfolio, and earlier held leadership roles at SoundCloud, Viacom Media Networks, Warner Bros., and Time Warner. Faye also advises companies and founders on scaling growth and navigating transformation, was named to Gold House’s 2024 A100 List. She holds degrees from the Wharton School and the University of Pennsylvania, as well as an MBA from Harvard Business School.

Ramayya Krishnan
Dean and W.W. Cooper and Ruth F. Cooper Professor of Management Science and Information Systems, Heinz College, Carnegie Mellon UniversityMore
Ramayya Krishnan is the Dean Emeritus and W. W. Cooper and Ruth F. Cooper Professor of Management Science and Information Systems at the H. John Heinz III College. A faculty member at CMU since 1988, Krishnan was appointed Dean when the Heinz College was created in 2009. He served three terms as its dean from 2009-2025. He is presently serving as the Director of the CMU NiST cooperative research center (https://www.cmu.edu/aimsec/) on AI measurement and Engineering, a center he helped establish in 2024.
Krishnan was elected in 2017 to serve as the 25th President of INFORMS. He served on the National AI advisory committee to the President and the White House (https://www.ai.gov/naiac/) and chaired its AI Futures Committee between April 2022 and April 2025. He chaired the AI academic council of the DOD’s CDAO through Dec 2025 and currently serve as an AI ambassador of the UKs RAI council.
Krishnan’s research interests focus on data analytics and data driven decision making. His current interests are focused on the intersection of these topics with AI measurement, evaluation and engineering in consequence domains such as healthcare. An application of particular interest is on AI and its impact on workflows and the workforce.
Presenters
Loujaina Abdelwahed
Head of Economic Research, Revelio Labs"How do workers feel about AI in the workplace?"
Artificial intelligence is changing the nature of work at an unprecedented pace. Most research on the impact of AI on the labor market has focused on estimating exposure and analyzing the effect of AI on employment, worker displacement, and wages. Yet very little is known about how workers experience the deployment of AI at work. As AI tools become embedded in most office jobs, understanding employees’ perceptions is critical for organizations and policymakers in managing this transition effectively, especially as the public narrative around AI and work has grown increasingly polarized thanks to statements by executives and the media about replacing workers with AI.
In this paper, I study how the diffusion of AI has affected employee sentiment, and what factors explain variation in employee experience in companies with high and low AI adoption. Specifically, I examine whether roles with higher AI adoption show larger declines in workers’ satisfaction.
I use data on employee reviews submitted on Glassdoor between 2019 and 2025. The dataset covers millions of reviews across all occupations in the US. Glassdoor asks workers to submit numerical ratings of employers (on a scale of 5) on different dimensions of their workplace experience. Employees are also asked to write free text on the pros and cons of their jobs. The numerical ratings along with the text allow for numeric evaluation of the impact of AI, in addition to text analysis on the reviews.
Preliminary results indicate that high-adopting companies experienced sharper declines in sentiment by approximately 0.17-points more than the decline in sentiment in the low-adopting companies after the launch of ChatGPT, after controlling for the role distribution and the level of exposure to AI. The decline in sentiment is observed across all the categories that employees rank. The text analysis shows that skill development dominates positive AI reviews, while workload concerns are the primary driver of negative reviews.
This paper directly relates to equity concerns related to the impact of AI on the labor market. The diffusion of AI is not experienced equally across workers. While some employees encounter AI as an opportunity for skill development and career growth, others experience it primarily as an added stressor. Workers in companies that are adopting AI at a fast pace are experiencing sharper declines in satisfaction and greater anxiety at the workplace, raising important questions about who bears the costs of technological change.
Daehwan Ahn
Assistant Professor, University of Georgia, College of Family and Consumer Science"Item-Mean Surrogates: Why Richer Individual-Level Data Fail to Improve LLMs as Human Surrogates"
Recent studies suggest that large language models (LLMs) can stand in for human respondents, and this has led many researchers and companies to use them as surrogates in social science and in market research. The premise is that richer persona descriptions, built from more individual-level data, allow a model to substitute for a specific person. We test this premise at scale, and we caution that it does not hold in the settings where it is most often assumed.
We study four datasets covering more than 400,000 participants and more than 6,000 response items, including a preregistered 19-study digital-twin corpus. At the aggregate level, LLMs perform well. Their average responses closely track average human responses to the same items (Pearson r = 0.78). Importantly, this success largely reflects item-mean tracking rather than individual prediction. Once each item’s mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, alternative model variants, and fine-tuning do not close this gap, and the best-performing configuration still reaches only 4.44%.
Our results show that current LLM surrogates behave mainly as item-mean trackers, not individual simulators. For researchers and firms adopting generative AI as a measurement instrument, aggregate fit should be treated as the starting point for surrogate validation, not as evidence that an LLM can stand in for individual respondents. We therefore urge that surrogate claims be validated for respondent-specific prediction and for the shape of human response distributions before LLMs are used in place of people.
Alon Bergman
Assistant Professor, The Wharton School, University of Pennsylvania"Projected Population Health Impact of Autonomous AI Primary Care in US Physician-Shortage Counties"
More than 100 million Americans live in federally designated primary care Health Professional Shortage Areas, with HRSA projecting a national shortfall exceeding 70,000 physicians by 2038. Shortages are geographically concentrated in Southern and rural counties, and preventable mortality and ambulatory care–sensitive hospitalizations track the geography of shortage closely. At the same time, frontier large language models now match or exceed practicing physicians on diagnostic vignettes, objective structured clinical examinations, and most quality and safety axes in head-to-head evaluations. These capabilities align closely with the core cognitive tasks of primary care. The open question is what happens—clinically and economically—when such generative AI agents are deployed at population scale in the precise geographies where human supply is scarcest.
We construct a steady-state county-level simulation covering 3,106 U.S. counties using 2025 Area Health Resources File physician supply data, CDC WONDER mortality statistics, and published causal estimates of the primary care supply–outcome relationship. Counties with supply below the policy threshold of 66.7 primary care physicians per 100,000 population (the 1:1,500 ratio recommended by the HRSA Negotiated Rulemaking Committee) are designated as treated. In these 2,303 counties (74.1% of all counties, home to 141.86 million residents), the AI PCP is modeled as adding effective physician-equivalent supply equal to the product of the local supply gap and patient adoption rate τ.
By quantifying the population health returns to deploying agentic AI as a supply-side complement in markets where the human labor pipeline has failed to clear, this work provides a quantitative foundation for business-model experimentation, investment priorities, and the design of regulatory institutions for agentic AI in safety-critical settings. The findings position generative AI not as a marginal efficiency gain in one service sector but as a structural intervention in the spatial allocation of cognitive labor—with the magnitude of the social return contingent on the institutions built around it.
Hemant Bhargava
Distinguished Professor of Management, Suran Chair in Technology Management at UC Davis, and Director of Center for Analytics and Technology in Society (CATS), Graduate School of Management, University of California, Davis"The Structure of U.S. State AI Policy: Evidence from a Multi-Dimensional Bill Classification"
Between 2019 and 2026, U.S. state legislatures introduced more than 2,500 bills addressing artificial intelligence, producing one of the most active subnational technology policy environments in recent memory. For firms building and deploying AI, this fragmented and fast-moving landscape is not a background condition but a direct input to strategy. Uncoordinated state action imposes a compliance tax that falls hardest on smaller firms, while larger firms may adopt the most stringent state’s standard as a de facto national baseline, concentrating regulatory power in a few jurisdictions. Yet basic questions about this activity remain unanswered: what are states actually legislating about, how do their priorities differ, and what predicts which proposals become law. Existing tracking efforts, including the National Conference of State Legislatures database, rely on keyword inclusion and a flat 24-category schema. Both choices undermine analytical use. Roughly a third of the bills in the source database have no substantive connection to AI, and the categorization scheme conflates regulatory mechanisms with policy domains, obscuring legislative intent.
We address these gaps with a measurement instrument and an empirical analysis of what it surfaces. The instrument combines a seven-element taxonomy with a multi-agent large language model classifier. The taxonomy is organized around legislative rationale, asking of each bill who is being constrained, to protect whom, and against what kind of harm. Its seven dimensions span product safety and accountability, AI inputs and intellectual property, sectoral use and application, AI markets and competition, institutional frameworks and processes, AI advancement, and a cross-cutting existential and societal risk dimension. The classifier operationalizes this taxonomy through ten specialized agents arranged in a graph-of-graphs architecture: a keyword pre-filter removes bills with no AI-specific terminology, an extractor and summarizer condense each bill to its policy-substantive content, seven dimension analysts score the bill in parallel so that judgment on one dimension does not anchor the others, and a judge agent reviews the analysts collectively and triggers targeted revision when scores and justifications diverge. This agentic decomposition mirrors the analytical labor a human policy coder performs by hand.
Léonard Boussioux
Assistant Professor, Foster School of Business, University of Washington"One Tool, One Taste? How Vibe Coding Homogenizes Web Design Without Its Builders Noticing"
Vibe-coding platforms now let anyone turn a plain-language description into a fully designed, deployed website in minutes, and millions of small businesses are acquiring their visual identity this way. We ask what happens to differentiation when so many businesses draw on the same generative layer. In a lab experiment, dozens of builders each designed a website for a different real business using the same vibe-coding tool. Embedding the resulting sites with a vision transformer and tracking a subset of build sessions second-by-second alongside concurrent think-aloud narration, we find that a wide range of independent design briefs collapses into a small number of recurring visual archetypes — not because inputs were similar, but because every builder runs the same generate-and-check loop, much of what they intend never reaches the prompt, and originality is almost never the criterion by which an output gets accepted. Most strikingly, this convergence is invisible to its creators: felt authorship, satisfaction, and perceived quality show no relationship to actual originality. We term this authored ignorance, and close with concrete platform responses that could let builders see, and escape, the house style they are unknowingly building toward.
Hancheng Cao
Assistant Professor, Goizueta Business School, Emory University"Generative AI, Lean Curation, and the Redesign of Innovation Evaluation"
As generative AI makes it increasingly easy to produce proposals, analyses, and prototypes, organizations face a discernment gap: creation increasingly exceeds the capacity to evaluate, prioritize, and direct scarce attention. This talk introduces Lean Curation as a vision for redesigning innovation evaluation. Inspired by Lean Manufacturing, it shifts attention from reducing material waste in production to reducing cognitive waste in how organizations allocate attention, compare opportunities, and learn. Rather than treating evaluation as a narrow act of scoring or selection, Lean Curation reframes it as a problem of prioritization, coordination, learning, and strategic alignment.
I then draw on academic peer review as a setting for examining how evaluation might be redesigned with large language models. Most LLM-assisted review systems reproduce inherited workflows by evaluating papers independently and assigning absolute scores. Yet such ratings are often compressed, inconsistently calibrated, and difficult to compare across reviewers. Pairwise comparison offers an alternative: LLM agents compare two papers at a time and judge which better satisfies multidimensional criteria. These relative judgments are aggregated into a global ranking using models such as Bradley–Terry, which estimate each paper’s latent quality from sparse outcomes.
Evidence from machine learning conferences suggests that pairwise LLM evaluation is more discriminative than rating-based review and can approach human review in identifying impactful work. At the same time, differences across research areas, novelty, and institutional concentration reveal risks. Lean Curation therefore emphasizes redesign rather than replacement: using AI to build scalable, informative evaluation systems while preserving human oversight, diversity, and strategic judgment.
Mengjie (Magie) Cheng
Assistant Professor, McCombs School of Business, University of Texas at Austin"Balancing Engagement and Polarization: Multi-Objective Alignment of News Content Using LLMs"
We study how media firms can use LLMs to generate news content that aligns with multiple objectives–making content more engaging while maintaining a preferred editorial stance. Using news articles from The New York Times, we first show that more engaging human-written content tends to be more polarizing. Further, naively employing LLMs to generate more engaging content can increase polarization. This has important managerial and policy implications: using LLMs without built-in controls to limit slant can exacerbate news media polarization. We present a constructive solution based on the newly proposed Multi-Objective Direct Preference Optimization (MODPO) algorithm, which integrates Direct Preference Optimization with multi-objective optimization techniques. Using an open-source LLM, we develop a new language model that simultaneously makes content more engaging while maintaining a preferred editorial stance. Furthermore, drawing on theory-driven content strategies, we find that our model achieves this balance by leveraging content characteristics that are strongly associated with polarization but have a comparatively smaller impact on engagement, e.g., minimizing provocative language while allowing more balanced perspectives. Our approach and findings can also apply to other settings in which firms use LLMs for content creation to achieve multiple objectives, such as advertising and social media.
Mina Cho
PhD Student, Carlson School of Management, University of Minnesota"Large Language Models Explain Humans Better Than Humans Themselves: Using LLMs for Tacit Knowledge Transfer"
Tacit knowledge, or the “know-how” embedded in experience, is difficult to articulate, making its transfer a challenge in organizations. As a result, valuable expertise is often poorly documented and lost when experts leave. Large Language Models (LLMs) offer a unique opportunity to address these challenges by identifying latent patterns embedded in experts’ behaviors, such as their responses or actions, and expressing them in natural language that can be transferred to novices.
This study examines whether LLMs can externalize tacit knowledge (i.e., transform into explicit knowledge) from experts’ behaviors and whether such externalized knowledge supports downstream decision-making and transfer to novices. We first test three prompt variations to understand how to leverage LLMs for externalizing tacit knowledge. We then examine the utility of LLM-externalized tacit knowledge across two studies. Results show that LLM-externalized tacit knowledge improves decision quality and enables novices to approach expert-level performance, often outperforming knowledge articulated by human experts. Mechanism analyses and robustness checks further demonstrate that LLMs meaningfully learn and extract knowledge from expert conversations (rather than simply copying their “style”), and that findings generalize across models and retrieval methods. Our findings suggest that LLMs can help overcome human experts’ articulation bottleneck, providing empirical support for Polanyi’s Paradox – that we can know more than we can tell. Overall, our study highlights the potential of LLMs as scalable tools for knowledge externalization and transfer, offering practical implications for organizational training and decision support.
Yuting Deng
PhD Student, Columbia Business School, Columbia University"AI-Moderated Interviews for Market Research and Digital Twins Calibration"
AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer digital twins. Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static methods. In a pre-registered, between-subjects study (N=317; 139 AI-moderated, 154 static, 24 human-moderated) with industry partners, we compare AI-moderated, human-moderated, and static interviews. AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant’s own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than simpler demographics-based personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in thinking styles between twins and humans and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).
Sorouralsadat Fatemi
Assistant Professor, California State University, Monterey Bay, College of Business"The Dark Side of Memory: Personalization Features and Biased Moral Advice in AI Assistants"
Memory features that let AI assistants personalize to individual users are now standard in deployed systems and are being adopted in settings where they inform consequential decisions. Yet prevailing measures of LLM sycophancy are one-dimensional, asking whether a model became more agreeable, and cannot distinguish uniform drift toward a fixed answer from directional bending toward the user’s own position.
We address this gap with a direction-aware measure, selective adherence, and evidence, from 30,799 paired evaluations across four models and four moral-dilemma datasets, that user memory shifts LLM moral judgments predominantly by bending them toward the user’s revealed position rather than toward a fixed pole, the size of the bend depending on what the memory reveals. Each of 700 scenarios was evaluated under a no-memory baseline and ten memory conditions (five profile types × two memory representations), with responses classified by a three-judge panel (Fleiss’ κ = 0.96 on stance). The shift is roughly twice as directional as uniform, an asymmetry corroborated by a confirmatory memory-by-lean interaction. We further observe a content gradient: memory of a user’s preferences and decision-style is associated with the largest adherence, separating clearly from a no-effect control, while emotional, expertise, and political signals show smaller effects. A positive effect appears in all four models and all four datasets, though directional dominance over uniform drift holds in three of the four.
One tension we do not resolve: raw histories induce more adherence than distilled profiles, opposite to prior conversational findings. Because memory is operationalized as short, controlled synthetic histories rather than real cross-session memory, the results are controlled-setting evidence requiring replication with real histories before informing deployment. For information-systems research and practice, the findings suggest that sycophancy audits of memory-enabled assistants should be direction-aware, since one-dimensional checks would understate a user-specific advice bias.
Manuel Hoffmann
Assistant Professor, University of California, Irvine"AI Adoption and Productivity in Healthcare"
Artificial Intelligence (AI) has the potential to reshape work processes and productivity in the health care sector. It presents a unique opportunity to enhance the economic value of professional work by automating time–intensive peripheral tasks, allowing highly skilled practitioners to focus on their core technical expertise. We investigate the adoption, productivity, and work transformation effects of AI scribes — ambient generative AI systems that document clinical conversations — within a large academic health network.
Leveraging a staggered rollout and granular data both from electronic health records and the AI scribe platform, we find that adoption of AI scribes is remarkably rapid, reaching near-universal coverage among eligible physicians within approximately a year. Physicians increase their usage post-eligibility and accept the vast majority of AI-generated notes based on audio recordings of physician-patient encounters. Eligibility for the AI scribe raises physician productivity by approximately 20 percent per physician-month, driven both by a higher volume of appointments and a shift toward more complex encounters. By reducing the marginal cost of documentation and alleviating administrative frictions, AI scribes enable a significant reallocation of time toward patient–centered care. A back-of-the-envelope calculation implies a cost-adjusted value of these productivity gains on the order of $1.6 million within the network to $1 billion when scaled to the national physician workforce. Ultimately, this longitudinal analysis demonstrates that when AI tools are designed to address pressing professional burdens, they have the potential to effectively bridge the implementation gap and drive improvements in both the quality and efficiency of care delivery.
Kartik Hosanagar
John C. Hower Professor; Professor of Operations, Information and Decisions; Faculty Co-Director, Wharton Human-AI Research, The Wharton School, University of Pennsylvania"Frontier AI Performance across the Business Disciplines: A Case-grounded Benchmark of Knowledge Work and Analytical Reasoning"
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The “case method” form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.
Hongxian Huang
PhD Candidate, Stern Business School, New York University"Generative AI Search on Content Creation Platform: Evidence from a Large-Scale Field Experiment"
Generative AI Search is an emerging platform-owned search-interface design that places AI-generated content above conventional ranked search results. For eligible queries, it provides users with a synthesized natural-language answer before they inspect or click the underlying sources, shifting the search experience from a retrieval-first, search-and-click model toward an answer-first format. On content creation platforms, where creator-supplied content is itself the informational product users seek, this design may affect both how often users search and how they subsequently consume the underlying content. Using a large-scale randomized field experiment involving 100,000 users on a leading Chinese content creation platform, we examine these two behavioral margins jointly. Treatment-group users received AI-generated content above the conventional ranked user-generated content (UGC) list when a submitted query satisfied the platform’s triggering rule for objective or sufficiently standardized experience-based answers, whereas control-group users always saw only the conventional ranked list.
Our results reveal a scale–intensity trade-off in content consumption. Generative AI Search expands the extensive margin of information seeking: users search more frequently, on more days, and across a wider range of content categories and semantic topics. This expansion spans all six substantive search intents, including those less aligned with the platform’s triggering rule. On the intensive margin, however, users who search consume less creator-supplied content per session: they browse and click less UGC. Despite this within-session contraction, aggregate UGC browsing and clicking increase, and more search sessions initiate downstream content consumption, which refers to subsequent within-platform content-page clicks after a user opens an initial piece of UGC from the search-results page. An activity–intensity decomposition shows that the expansion in search activity more than offsets the decline in average content consumption per session. Thus, Generative AI Search can reduce the consumption of creator-supplied content within individual sessions while increasing its aggregate consumption, demonstrating why lower per-session consumption need not imply lower overall demand for creator-supplied content.
Minji Jang
Postdoctoral Research Associate, INSEAD Ethics & Social Responsibility Initiative"That’s What Claude Said: AI Overreliance as Unwarranted Closure of Inquiry"
When does reliance on AI become overreliance? Existing research largely treats overreliance as miscalibrated trust or error: actors overrely when they trust AI beyond its demonstrated capabilities or when they accept recommendations that turn out to be wrong. Trust-based accounts cannot deliver consistent diagnosis: ‘trust’ is a normatively loaded notion open to multiple interpretations, and it is unclear when it becomes excessive. Error-based accounts can deliver consistent diagnosis but fail to account for the agency of the user and the complex layers of user-machine relationships, such as vulnerability and accountability.
We instead propose to understand AI overreliance as unwarranted closure of inquiry. Drawing on Friedman (2019, 2020), we define closure as settling a question, distinct from trust and reliance. Closure itself is not objectionable: the search must stop somewhere, and AI can help close the right questions. Defining overreliance as closure without warrant shifts the governing question from “Do users trust AI too much?” to “When are actors warranted in treating AI outputs as settling questions before them?”
We develop a three-layer framework focusing on what is being closed, why actors close, and when closure is warranted. This framework provides conceptual clarity on when seemingly excessive dependence on AI is warranted, and when seemingly successful dependence might not be. It further delineates the mechanism by which overreliance incurs organizational costs. For managers, it shifts the focus from naïve or excessive policing of AI to governing closure.
Siyuan Bruce Jin
PhD Candidate, HKUST Business School and The Wharton School, University of Pennsylvania"Codified Expertise: How Generative AI Changes Temporal Coordination in Distributed Teams"
We study whether generative AI complements or substitutes for synchronous coordination in distributed software work. Using the staggered adoption of an AI coding assistant by approximately 76,000 developers across 7,074 teams at a global technology firm during 2024–2025, we show that the productivity gain from AI adoption is largest when synchronous expert input is hardest to obtain. AI adoption increases quality-gated releases by 9.7% on average, and this gain rises systematically with a team’s mean pairwise time-zone separation. A decomposition of the dispersion measure shows that the binding relationship is not the average pairwise distance in the team, but the temporal distance between non-senior developers and the senior technical gatekeeper. Developers farther from the senior gain contribution share, move upward in the within-team contribution ranking, and exhibit less week-to-week volatility in output after adoption. We find no comparable moderation by physical distance, indicating that the friction AI relieves is temporal rather than spatial. Whereas the value of earlier coordination technologies increased with greater collaborative hours, AI substitutes for synchronous expert input precisely where those hours are scarce. The evidence identifies generative AI as a codification technology. It embeds expert routines and makes them available when synchronous expertise is unavailable.
Enoch Hyunwook Kang
PhD Student, University of Washington"Personalized AI Alignment Revisited: The Necessity and Sufficiency of User Diversity"
Personalized AI alignment, i.e., adapting AI systems to individual preferences, is becoming central as AI agents increasingly recommend, generate, negotiate, and act on users’ behalf. Yet empirical evidence remains mixed: personalized AI sometimes improves over non-personalized baselines, but often does not. In this paper, we provide an if-and-only-if characterization of when personalized AI alignment works. The key quantity we identify is decision-relevant user diversity: the extent to which the feedback population covers the preference distinctions that can change personalized decisions. When this diversity is present, harmful reward alternatives are statistically separated from the truth; when it is weak, a model can fit observed feedback while still making wrong decisions for users whose preferences depend on poorly measured distinctions. Specifically, we show that decision-relevant user diversity is necessary and sufficient for bounded online decision regret, and that its offline analogue yields logarithmic sample complexity.
This characterization turns personalization into a platform problem of managing the feedback population. In stable environments, paying users may eventually reveal the preference distinctions needed to personalize well for them, but in realistic agent settings, prompts, products, tasks, inventories, campaigns, and user needs change over time, so paying users may reveal the relevant distinctions too slowly for the current decision window. The platform may therefore need to acquire costly feedback from free users whose preferences expose directions that are weakly measured among paying users but important for serving them.
We formalize this problem as fixed-budget free-user acquisition for paying-user personalization. Because the value of free users is initially unknown, we propose an optimistic acquisition rule that learns which free users are informative while spending the acquisition budget, and prove a near-optimal fixed-budget guarantee. Synthetic simulations and PRISM language-alignment experiments show that personalized alignment depends not only on the learning algorithm, but also on how the platform manages the preference feedback population.
Arvind Karunakaran
Assistant Professor of Management Science & Engineering, and Sociology, Stanford University"Experimentalist Intensification Governance: Managing Worker Negative Consequences Associated with Generative AI Experimentalist Work"
How does work intensification occur in organizations around the use of emerging technologies such as Generative AI (GenAI)? And how might such work intensification be mitigated? Existing studies examine the work intensification that occurs when workers use GenAI for individual productivity, but overlook the work intensification that may occur when workers use these technologies to create organizational solutions (solutions that have the potential to create value for a broad set of organizational stakeholders). In our comparative, qualitative field study of medical and legal workers using GenAI to create organizational solutions in an academic medical center and law firm, we find that the use of GenAI resulted in new kinds of work intensification related to workers’ engagement in experimentalist practices such as distributed trialing, collective review and revision of results, and aligning and hustling. More importantly, we find that the negative consequences for workers associated with experimentalist practices varied according to different forms of what we call experimentalist intensification governance structures in the two organizations. This study advances the sociological literature on organizational control, invisible work, and voluntary work intensification by demonstrating how experimentalist intensification governance can mitigate negative worker consequences of experimentalist work around emerging technologies such as GenAI.
Riitta Katila
W. M. Keck Sr. Professor of Management Science, Faculty Director of the Stanford Technology Ventures Program and HAI Sabbatical Scholar at Stanford Institute for Human-Centered Artificial Intelligence, Stanford University"Selecting the Future: When AI Outperforms Experts and When Experts Outperform AI"
Organizations generate more ideas than they can fund, making selection central to the variation–selection–retention (VSR) process. We ask whether large language models can substitute for, or complement, human experts in selecting early-stage technology ventures in ways that predict downstream outcomes—and when each performs best. We study a Beijing province-level startup competition with 981 startups. Each venture was independently scored by five human experts. In parallel, we generated five LLM evaluations per startup using expert-like personas. We link scores to three-year outcomes capturing innovation potential (patent applications) and growth potential (open job positions). Human evaluations strongly predict subsequent innovation, while AI evaluations more reliably predict hiring. These differences are systematically moderated by text characteristics of the business plan: lexical complexity strengthens AI’s predictive validity, whereas abstractness tends to advantage humans, supporting the underlying mechanisms.
Hyunso Kim
Research Fellow, SNU Business School, Seoul National University"Generative AI Fuels Solo Entrepreneurship, but Teams Still Lead at the Top"
Recent advances in generative artificial intelligence (AI) are reshaping who enters entrepreneurship, but not who reaches the top of the quality distribution. Using data on over 160,000 product launches on Product Hunt, we find that entrepreneurial entry increased sharply following the public release of ChatGPT-3.5, driven disproportionately by solo entrepreneurs. This shift toward solo entry is particularly pronounced in categories that historically favored team-based ventures. However, much of this growth reflects low-commitment, experimental entry and does not translate into greater representation among the highest-quality outcomes. Team-based ventures are increasingly dominant in the top tiers of platform rankings. These findings reveal a growing divergence between the quantity and quality of entrepreneurial activity, and a shift in the binding constraint on entrepreneurship from venture initiation to sustained development.
Leonard Kinzinger
PhD Candidate, Technical University Munich (TUM)"Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?"
LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts. Neither setup speaks to the operationally most relevant case for marketing practice: building detailed individual twins from the pre-existing heterogeneous panel data that firms already accumulate through CRM systems, loyalty programs, and repeat surveys. We construct detailed individual-level twins from the German Socio-Economic Panel (SOEP) and evaluate them across a 3×5×2×2 construction-method grid that covers three open-weights LLMs, five cumulative information depths ranked by normalized Shannon entropy, two embedding methods, and two reasoning modes, scoring over 2.1 million twin responses on 500 participants and 183 held-out questions. Twin quality rises with information depth but with diminishing returns past the 75 percent entropy quartile, which acts as a cost-efficient Pareto point relative to the best-performing 100 percent cells. Switching the embedding from a narrative persona summary to a raw dialog history of past responses raises hold-out accuracy in every model-by-reasoning cell at the 100 percent depth, while an explicit thinking mode raises rank-order correlation without moving accuracy. Best-cell accuracy reaches 78.8 percent and Fisher-z correlation reaches r=0.590 on the SOEP held-out evaluation set. The findings suggest that twin-based market research is no longer gated by data design, but by item volume, model selection, and a small set of construction-level decisions that this paper now maps.
Adam Kuzee
Pre-doctoral Researcher, MIT FutureTech, Massachusetts Institute of Technology"Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks"
We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and (ii) rising tides where the increase in AI capabilities is more continuous and broad-based. We test for these effects in preliminary evidence from an ongoing evaluation of AI capabilities across over 3,000 broad-based tasks derived from the U.S. Department of Labor O*NET categorization that are text-based and thus LLM-addressable. Based on more than 17,000 evaluations by workers from these jobs, we find little evidence of crashing waves (in contrast to recent work by METR), but substantial evidence that rising tides are the primary form of AI automation. AI performance is high and improving rapidly across a wide range of tasks. We estimate that, in 2024-Q2, AI models successfully complete tasks that take humans approximately 3-4 hours with about a 50% success rate, increasing to about 65% by 2025-Q3. If recent trends in AI capability growth persist, this pace of AI improvement implies that LLMs will be able to complete most text-related tasks with success rates of, on average, 80%–95% by 2029 at a minimally sufficient quality level. Achieving near-perfect success rates at this quality level or comparable success rates at superior quality would require several additional years. These AI capability improvements would impact the economy and labor market as organizations adopt AI, which could have a substantially longer timeline.
Recent evidence by Kwa et al. (2025) suggests that, as models improve, AI capabilities surge abruptly for tasks that previously appeared out of reach, as if a “crashing wave” suddenly reaches them (our characterization). In this paper, we contrast this crashing wave phenomenon with a “rising tide,”, in which performance is lifted more broadly across the task space. The central difference between the two phenomena is the slope of the relationship between AI success on tasks and (log) task duration. For crashing waves, this relationship can be well described by a steep logistic curve. AI progress is then a rightward shift of the curve, which translates into large, concentrated automation for tasks near the tipping point due to sudden improvements in capabilities of what systems can do. In practice, this would lead to harsh surprises for human workers. Over just a short period of time, they would observe AI models going from nearly always failing to nearly always succeeding. By contrast, rising-tide automation has a flatter success–duration relationship, with AI performance being more similar across tasks of different lengths. This would still be represented by a logistic curve, but a much flatter one. The same amount of AI progress would then translate into more gradual automation under the rising-tide view, such that individual workers are less likely to be blindsided by AI. A rising tide could, however, still be quite disruptive if it happens quickly. The main insight of this paper is that, across a large set of realistic and representative labor-market tasks addressable by LLMs, the downward slope between task success and task duration is, on average, surprisingly flat — i.e., more consistent with a rising tide rather than a crashing wave. Our analysis draws upon a broader ongoing research effort that collects novel evaluations of LLM outputs by domain experts across more than 40 models and covering over 20,000 unique task examples (“instances”) based on more than 10,000 ONET tasks that are at least partially text-based. Outputs are scored by human evaluators with relevant on-the-job experience. Given the paper’s primary focus on automation, we center our analysis on success measures defined by expert evaluations indicating that the LLM output required no human intervention to be at least minimally successful.
Dokyun DK Lee
Professor of Information Systems and Computing & Data Science, Boston University"Counting the Species of Ideas: A Bayesian Ecological Survey of LLM Idea Spaces"
Large language models (LLMs) increasingly generate product concepts, research questions, and solutions. Widespread LLM use may create algorithmic monoculture as users converge on the same machine-generated ideas. Measuring this risk requires mapping the full creative idea space a model can access, including unseen or rare ideas, rather than merely counting unique responses in a finite sample. We present a Bayesian ecological framework for surveying an LLM idea space. We treat distinct ideas as species, generations as captures, and repetitions as recaptures. This permits heterogeneous capture-recapture estimation under assumptions suited to repeated, stateless LLM interaction. For a fixed model, task, prompt, and decoding policy, we characterize the idea space through N, p, and v. N is the number of ideas a model can access, p is the probability of generating each idea, and v is the value of each idea for a given task.
Across diverse open-source and proprietary LLMs, we show that this approach maps idea spaces more accurately and efficiently than brute-force sampling while requiring a fraction of the computational time. It also estimates the shape of p, separating a dominant core of recurring ideas from a long tail of rarely surfaced possibilities. We then compare models and LLM creative elicitation strategies from recent computer science research. The framework provides a practical first step toward diagnosing and correcting algorithmic monoculture at scale. By enabling accurate and efficient measurement of LLM idea spaces, this framework provides a crucial foundation for developing prescriptive interventions that counter LLM-induced creative homogenization across organizations.
J. Frank Li
Assistant Professor, UBC Sauder School of Business"Generative Algorithms as General Algorithms"
Generative artificial intelligence (GenAI) is transforming digital platforms from retrieval-based information systems into interactive systems that synthesize information, explain alternatives, and support user sense-making. This study examines whether generative algorithms function as general algorithms that subsume established platform functions such as recommendation and search. We combine China’s mandatory algorithm filing records with monthly active user data from the Apple App Store and a demand-side survey of 517 users of both Weibo and Taobao. The filing data show that generative algorithms diffuse across all 23 App Store categories and concentrate in productivity-oriented domains, whereas conventional recommendation and search algorithms remain concentrated in entertainment- and information-oriented domains. Generative algorithms also rapidly replace conventional algorithms in search and recommendation filings. Difference-in-differences estimates indicate that generative recommendation and generative search filings increase monthly active users by 27.2 percent and 38.8 percent, respectively, while conventional algorithms have no detectable effect. Competitor-based spillover analyses and heterogeneity tests further support the interpretation that generative algorithms improve user-facing intermediation, especially by supporting both demand fulfillment and demand formation. Survey evidence from Weibo and Taobao users complements the app-level results: users who use or prefer GenAI search and recommendation features open the corresponding apps more frequently, and instrumental-variable estimates using AI chatbot usage support this demand-side mechanism. The study contributes to information systems research by theorizing generative algorithms as cognitive intermediation technologies with general-purpose properties in the digital platform economy.
Xinyu Liang
Assistant Professor, INSEAD"Does AI Crowd Out Context? Evidence from Clinical Radiology"
Human-AI collaboration is often deployed on the premise that algorithmic precision and human contextual knowledge are complementary. We show a limit to this premise: displaying AI changes how experts use the contextual soft information they alone observe. In an experiment in which radiologists evaluate chest X-ray cases under independently varied access to AI predictions and clinical history, performance without AI tracks the value of clinical history: errors fall most on the patient–pathology pairs where history is most informative. Displaying AI alongside that history reverses the relationship, and the combined workflow performs relatively worse precisely on the pairs where history has the most to offer. A Bayesian signal-extraction framework separates two explanations. Because AI sharpens the image-based channel, a rational decision maker should rely less on clinical history even absent behavioral distortion; yet the observed attenuation is roughly 2.6 times this rational benchmark. Clickstream data are consistent with attentional redirection rather than reduced effort: total effort does not fall, but active engagement with clinical detail does when AI is on screen. The performance cost, measured as agreement with a fully informed subspecialist benchmark, is concentrated where it matters most: the combined workflow underperforms clinical history alone on 34.5% of patient–pathology pairs, rising to 59% among the highest-value pairs. The realized value of human contextual knowledge is thus endogenous to the AI workflow: effective deployment must account not only for average algorithmic accuracy, but also for whether the workflow lets algorithmic advice and human contextual judgment complement rather than displace each other.
Benjamin Lira
Postdoctoral Scholar, The Wharton School, University of Pennsylvania"Synthetic Contact with AI Reduces Cross-Partisan Animosity"
Americans’ warmth toward members of the opposing political party has fallen sharply over the past three decades — yet meaningful cross-partisan contact remains scarce, in part because people actively avoid it. Across five preregistered studies (total N = 3,960 U.S. partisans), we test whether brief conversations with AI chatbots representing the political outgroup can substitute for the contact people shun. Synthetic contact first lowers the barrier to entry: partisans would endure almost twice as long contemplating their own mortality to avoid a human outgroup partner as an AI one. These conversations then correct the misperceptions that fuel division. At baseline, Democrats placed Republicans more than a standard deviation past their actual position on environmental consumption attitudes — enough to flip the average Republican from supportive to opposed — and a single ten-minute conversation with an outgroup chatbot corrected those beliefs and warmed affect in a within-person study of both parties. A three-arm experiment ruled out pure engagement and sociality as drivers. Synthetic contact also moved behavior, in a sample of both parties and on a more affectively charged issue: participants who spoke with an outgroup bot about immigration were six percentage points more likely than controls to choose to have a real conversation with a partisan from the other side. A final study tested whether these gains last: the warmth effect replicated immediately in a new sample; most of it faded within a week, with a small residual concentrated among the most extreme partisans. Analyzing conversation content showed that information, more than friendliness, distinguishes outgroup bots from control chatbots. Together, these findings establish synthetic contact as a scalable, behaviorally consequential, and — unlike face-to-face contact — widely acceptable form of cross-partisan engagement.
Chaoran Liu
Assistant Professor, Peking University (UK Campus)"Generative AI and Content Homogenization: The Case of Digital Marketing"
Entrepreneurs and small businesses often face resource constraints that limit their ability to produce high-quality marketing content, hindering business growth. Generative artificial intelligence (Gen AI), particularly large language models such as ChatGPT, offers an affordable and scalable solution for content creation. However, widespread adoption of these tools may produce unintended consequences, including content homogenisation, where AI-generated content becomes increasingly similar across firms. Such homogenisation may reduce brand distinctiveness and weaken consumer engagement over time.
This study examines the impact of Gen AI on content homogenisation and marketing performance in the restaurant industry, a highly fragmented sector in which approximately 70% of establishments are independently owned. We exploit Italy’s nationwide ChatGPT ban in April 2023 as a natural experiment to identify the causal effects of restricted Gen AI access on Instagram marketing content. Using a difference-in-differences framework, we compare restaurants in Milan before, during, and after the ban to assess changes in content similarity, posting behaviour, and consumer engagement.
The findings show that restricting access to ChatGPT significantly reduced content similarity across multiple linguistic dimensions. During the ban, restaurants in the treatment group experienced relative decreases of 15% in lexical similarity, 12% in syntactic similarity, 2% in semantic similarity, and 3% in language-style similarity, indicating that access to Gen AI contributes to greater content homogenisation. At the same time, the ban reduced posting frequency and post length but increased consumer engagement, with average Instagram likes rising by approximately 3.5%. Experimental evidence confirms the underlying mechanism by demonstrating that increased content homogenisation causally reduces consumer engagement. These results suggest that while Gen AI improves the efficiency of marketing content production, excessive reliance on AI-generated content may diminish originality and consumer appeal. By highlighting an important unintended consequence of Gen AI adoption, this study provides valuable insights for entrepreneurs and small businesses seeking to balance the productivity benefits of AI with the need to maintain authentic, distinctive, and engaging brand communication.
Lan Luo
Assistant Professor of Marketing, Yale School of Management"How Visual Designs Drive Success: Interpretable Generative AI for Data-Driven Design"
Visual designs are often used in marketing (e.g., packaging, ads, media covers) to achieve a variety of business outcomes, like improved sales, click-through rates, and brand attitudes. Since designs are complex, unstructured data, it is difficult to determine what features drive their success in a way that is interpretable and managerially actionable. To address this challenge, I develop a novel methodological framework to automatically discover what interpretable features make visual designs in a given domain successful. I first leverage a deep generative text-to-image AI model (by fine-tuning Stable Diffusion 3.5 in my application) that adopts the role of designer and enables visual designs to be described by low-dimensional design representations. Then, I apply a novel adaptation of cutting-edge “mechanistic interpretability” methods—specifically “sparse autoencoders” typically applied to large language models—to scalably discover a taxonomy of interpretable and managerially relevant features predictive of success from these design representations. Finally, I generate image redesigns by manipulating features of interest to help managers scalably pilot data-driven design changes. I apply this framework to discover how book cover redesigns predict sales on Amazon.com using a unique dataset I collected of over 160,000 books. I discover a diverse set of interpretable features related to illustration, typography, composition, and layout. I then create realistic cover redesigns predicted to improve sales by manipulating those features (e.g., redesigns with lower contrast and less separation of text and graphical elements). In a holdout analysis with a rich set of control variables, including just 30 of these discovered features (out of 9,728) improves variation explained in sales by nearly as much as prices and by more than reviews. Back-of-the-envelope calculations suggest that a large publisher could leverage this subset of features to increase annual revenue for the whole publisher by over $9.1 million, reflecting a change in sales equivalent to introducing an 8.5% price discount. In a lab study, I find causal evidence that the proposed methodological framework can redesign covers to significantly improve preferences, and that generative AI can help level the playing field in the publishing industry.
Shilei Luo
PhD Student, Washington University in St. Louis"Behavioral Transfer in AI Agents: Evidence and Privacy Implications"
Artificial intelligence agents are increasingly deployed to act on behalf of humans in digital environments. Whether agents accumulate owner-specific context, with consequences for both their behavior and their owners’ privacy, is unknown. On a platform where agents that owners use for routine tasks act autonomously, we compare 10,659 agents to their owners’ independently recorded social media activity. Agents systematically reflect their owners across topics, values, affect, and style. Transfer persists for agents without explicit configuration; pairs aligned on one dimension align on others, consistent with accumulated owner-agent interaction during ordinary use. Notably, 34.6 percent of agents publicly surface personal information about their owners absent from their configuration, and disclosure likelihood rises with behavioral transfer. These findings reveal a previously undocumented privacy pathway with governance implications.
Xueming Luo
Charles Gilliland Distinguished Chair Professor of Marketing, Professor of Management Information Systems, Temple University, Fox School of Business"Can Empathy Lead to Sales in Influencer Marketing? Multimodal AI Learning of TikTok Videos"
Influencer marketing has become an important part of marketing strategy. But what makes some influencer content more likely to drive sales? This paper examines the role of perceived empathy. While empathy is often assumed to boost sales, a multimethod investigation examines whether that is actually the case. First, we analyze purchases from almost 4,000 influencer livestreaming TikTok videos. To do so, we train an interpretable multimodal graph deep learning model that integrates interactions among audio, linguistic, and visual information to most accurately predict perceived empathy. Second, to directly test causality, we conduct pre-registered experiments, including a randomized field experiment among 6,000 customers. Results consistently find an inverted U: While some perceived empathy can boost sales, too much can backfire (because it suggests the influencer may just be pretending to be empathetic to push products). Overall, this work offers insights to improve influencer marketing, provides a tool for managers to benchmark and select influencers, and empowers researchers with an interpretable multimodal deep learning method to concurrently examine audio, textual, and visual cues from video data and how they interact to shape marketing outcomes.
Chengfeng Mao
PhD Candidate, Sloan School of Management, Massachusetts Institute of Technology"The Validation Dilemma of LLM Survey Surrogates"
Researchers propose large language model (LLM) surrogates as proxies for human respondents in surveys and behavioral experiments. We consider two regimes of surrogate deployment on a target item: before any human responses are collected (R1) and after a small calibration sample is collected (R2). We argue that its item-specific accuracy is unverifiable under R1 and that it adds no demonstrated value under R2. In R1, no reviewed evaluation specifies how performance on previously observed items predicts error on a new item; verification therefore requires collecting responses to that item. In R2, the surrogate must improve inference beyond the collected responses and a simple model fitted to the same data. None of the 56 papers we code provides the required evidence in either setting. We test these requirements on Twin-2K-500 and SocSci210, where surrogates predict responses from structured respondent profiles, evaluating four item-level summary statistics. Under R1, errors vary widely across items, and no tested method reliably identifies in advance the items on which the surrogate will be accurate. In R2, combining LLM surrogates with prediction-powered inference (PPI++) reduces variance by at most a few percent; replacing the LLM with regression yields similar performance. This creates a validation dilemma: surrogates are unverifiable without target-item responses and largely redundant once calibration responses exist. We state four falsifiable conditions that would overturn these conclusions.
Alex Mari
Assistant Professor of Marketing, American University of Bahrain; University of Zurich"The Right Time for AI Empathy in Voice Commerce: Attribution Processes and Product Category Shape Decision Satisfaction"
When should an AI shopping assistant express empathy? As generative voice assistants become increasingly embedded in retail journeys, firms must determine not only whether these agents should display empathy, but also how their empathic responsiveness should adapt across shopping contexts. Drawing on attribution theory, this research examines how AI empathy shapes consumer decision satisfaction in voice commerce and identifies the psychological mechanisms and boundary conditions that determine its effectiveness.
Across three experiments involving 748 participants, consumers searched for, compared, and selected products from the live eBay catalog using an empathic voice interface (EVI) powered by the large language model Claude. The findings show that AI empathy enhances decision satisfaction both directly and indirectly. Empathic voice assistants are perceived as more transparent and less strategically manipulative, thereby improving consumers’ evaluations of the decision-making process.
However, empathy does not operate uniformly across product categories. For hedonic purchases, such as scented candles, empathy has a stronger and more immediate effect on decision satisfaction. For utilitarian purchases, such as batteries, consumers respond more cognitively: empathy improves satisfaction primarily by increasing perceived transparency and reducing concerns about manipulative intent.
Building on these findings, the final study explores a context-adaptive model in which the assistant dynamically calibrates its level of empathy to the consumer’s purchase goal in real time. Preliminary evidence suggests that pairing stronger empathy with hedonic shopping and more restrained empathy with utilitarian shopping may improve decision outcomes.
This research challenges the assumption that more AI empathy is always better. Instead, it positions empathy as a dynamic design variable that should be calibrated to the purchase context in real time.
Abhishek Nagaraj
Associate Professor, Haas School of Business, UC Berkeley"CentaurBench: Benchmarking LLM Augmentation on Occupational Tasks"
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent’s performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model’s guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Jiaxin Pei
Assistant Professor, School of Computing, UT-Austin"How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks"
The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models are more token-efficient? and (3) Can agents predict their token usage before task execution? In this paper, we present the first systematic study of token consumption patterns in agentic coding tasks. We analyze trajectories from eight frontier LLMs on SWE-bench Verified and evaluate models’ ability to predict their own token costs before task execution. We find that: (1) agentic tasks are uniquely expensive, consuming 1000x more tokens than code reasoning and code chat, with input tokens rather than output tokens driving the overall cost; (2) token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30x in total tokens, and higher token usage does not translate into higher accuracy; instead, accuracy often peaks at intermediate cost and saturates at higher costs; (3) models vary substantially in token efficiency: on the same tasks, Kimi-K2 and Claude-Sonnet-4.5, on average, consume over 1.5 million more tokens than GPT-5; (4) task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend; and (5) frontier models fail to accurately predict their own token usage (with weak-to-moderate correlations, up to 0.39) and systematically underestimate real token costs. Our study offers new insights into the economics of AI agents and can inspire future research in this direction.
Lily Poursoltan
PhD Candidate and Senior Data Scientist, University of California San Diego"Designing Human-Agent Ecosystems: Interaction-Level Fit and Selective Delegation to Autonomous Agents"
As organizations embed generative AI into professional workflows, a central question is how work should be allocated between AI agents and human professionals: which tasks or task components should be delegated, where human involvement should remain, and how responsibility should be distributed across the workflow. We examine this question through task-technology fit, arguing that human-in-the-loop generative AI requires attention not only to task-type fit but also to interaction-level fit, or whether a generated contribution meaningfully reduces human effort.
We study AI-assisted patient portal messaging at an academic medical center using thirteen months of data covering 12,030 clinician-message interactions, 1,017 clinicians, and more than thirty outpatient specialties. Because clinicians choose whether to initiate from an AI-generated draft, we use a within-clinician design and a doubly robust double machine learning estimator to flexibly adjust for high-dimensional features capturing message semantics and complexity, clinician workload, and temporal context.
AI-draft initiation is associated with a 20% reduction in clinician response time after accounting for discretionary use. However, benefits vary across tasks. Time savings are greatest for standardized requests, including medication refills and paperwork, and decline as interpretive demands increase. Laboratory-result questions show no reliable improvement because the system lacks access to result values and case-specific clinical context, identifying information architecture as a boundary condition on agent performance.
We further develop an ex ante fit score that supports selective delegation by identifying interactions in which drafts are most likely to reduce effort. The findings reconceptualize delegation as an interaction-level allocation problem and show how organizations can effectively distribute work across humans, general-purpose agents, and specialized agents according to predicted fit, information requirements, and professional judgment.
Ruben R Salas
PhD Student, The Wharton School, University of Pennsylvania"Tell Me the Truth: Conjoint Analysis in the LLM Era"
Large Language Models now mediate consumer choice on both sides of the market. Marketing researchers use them as silicon respondents in conjoint panels, and consumers read AI-generated rationales. Both uses rest on the assumption that the model’s stated reasoning matches the attribute information actually driving its choice. Conjoint analysis has long shown that the same assumption fails in human respondents. What people say drove their choice diverges from what actually drove it. We test whether LLM respondents inherit the same pattern. Open-weight models let us measure both signals on the same product pair. We score the stated signal from chain-of-thought mention shares and the revealed signal from integrated-gradient attributions on the choice probability. Across seven open-weight model families, the two disagree. A natural cause to explore is reinforcement learning from human feedback, which aligns the verbal layer to human-stated preferences that themselves carry the gap. The data tell the opposite story. On average, RLHF mildly reduces the gap, with heterogeneity across families. We then ask whether the divergence matters for the consumer. Persuasion-knowledge theory predicts that surfacing a socially loaded driver triggers skepticism and lowers uptake. We pre-registered an experiment contrasting an AI shopping assistant’s natural rationale with an honest rewrite of it. Honest disclosure raises consumer uptake relative to the natural rationale, against the theoretical prediction. The gain is associated with higher perceived competence.
Neha Sharma
Assistant Professor, The Wharton School, University of Pennsylvania"Generative AI Shifts Technical Knowledge Production Toward Recombinant Novelty"
Public debate about generative AI often centers on whether LLMs will replace humans or instead shift work toward new forms of human–AI complementarity. We study how this shift unfolds by examining what questions people continue to ask other humans even after the broad availability of GPT-style tools. Using longitudinal data from a large technical Q\&A community for software developers, we analyze the content and sources of questions posted after the `LLM shock’ (late 2022). We find that while overall question volume declines, the questions that persist are disproportionately novel—but importantly, this novelty is qualitatively different from novelty in prior years. Post-LLM novel questions are less about entirely new technical domains and more about recombining existing knowledge domains, and they are more likely to involve niche tags rather than established `popular’ topics. This can also be seen in the hollowing of the core and connection to the frontier of knowledge network structure. Notably, this shift is driven by both selection and migration: new entrants are more likely to ask novel questions, and incumbent users become increasingly likely to ask these recombinant, niche questions over time, consistent with within-user reallocation of routine queries to LLMs. Together, these patterns clarify what kinds of knowledge work remain human-mediated after GPT and highlight implications for the future of open knowledge communities—including the sustainability of the data pipeline that supports further LLM development.
Jesse Silbert
Postdoctoral Research Associate, MIT FutureTech, Massachusetts Institute of Technology"Making Talk Cheap: Generative AI and Labor Market Signaling"
Large language models (LLMs) like ChatGPT have significantly lowered the cost of producing written content. This paper studies how LLMs, through lowering writing costs, disrupt markets that traditionally relied on writing as a costly signal of quality (e.g., job applications, college essays). Using data from Freelancer.com, a major digital labor platform, we explore the effects of LLMs’ disruption of labor market signaling on equilibrium market outcomes. We develop a novel LLM-based measure to quantify the extent to which an application is tailored to a given job posting. Taking the measure to the data, we find that employers had a high willingness to pay for workers with more customized applications in the period before LLMs were introduced, but not after. To isolate and quantify the effect of LLMs’ disruption of signaling on equilibrium outcomes, we develop and estimate a structural model of labor market signaling, in which workers invest costly effort to produce noisy signals that predict their ability in equilibrium. We use the estimated model to simulate a counterfactual equilibrium in which LLMs render written applications useless in signaling workers’ ability. Without costly signaling, employers are less able to identify high-ability workers, causing the market to become significantly less meritocratic: compared to the pre-LLM equilibrium, workers in the top quintile of the ability distribution are hired 19% less often, workers in the bottom quintile are hired 14% more often.
Fangchen Song
PhD Candidate, University of Texas at Austin"Better AI, Less Human Value Added? Evidence from a Large-Scale Field Experiment with Early-Career Knowledge Workers"
As AI models advance rapidly, organizations face a fundamental question: How can humans continue to create value beyond increasingly capable machines? We investigate this question using a field experiment involving 523 early-career professionals at a large global professional services firm who completed knowledge-intensive tasks while interacting with AI agents built on two models, GPT-3.5 Turbo and GPT-5.1. We evaluate human contribution relative to an AI-only baseline, which allows us to isolate the incremental value created — or destroyed — by humans beyond what AI can accomplish on its own.
Our analyses reveal several key findings. First, the AI-only baseline increases dramatically, rising by 83% between the two AI regimes. Second, the composition of human-AI collaboration profiles changes substantially: the proportion of individuals who outperform AI (“AI amplifiers”) declines sharply, while the proportion who underperform the AI baseline (“AI apprentices”) increases considerably. Third, there is a reduction in human value-added as AI capability advances. Even AI amplifiers generate less incremental value beyond AI, whereas AI apprentices destroy substantially more value relative to the AI baseline. Model-based analyses controlling for task characteristics, prior AI experience, training, tenure, educational background, and collaboration modes produce the same conclusion: although stronger AI improves overall human-AI performance, it simultaneously reduces the marginal contribution attributable to humans.
We investigate two plausible mechanisms underlying this phenomenon. First, better AI models frequently generate strong initial solutions, but because generative AI models are highly responsive to user instructions, subsequent human interventions may unintentionally redirect interactions away from these high-quality starting points by introducing noise, emphasizing less relevant aspects of the task, or moving the conversation away from an optimal solution path. Second, as AI becomes more capable, some individuals may reduce their involvement in directing and managing the problem-solving process, relying excessively on AI despite the continued need for human judgment, verification, and integration.
Taken together, our findings suggest a shift in value creation in human-AI collaboration with advances in AI models. Organizations face a new challenge of how to train users who can preserve human agency and continue to create value in the presence of increasingly capable machines.
K. Sudhir
James L. Frank Professor of Private Enterprise and Management, Yale University"The Production-to-Orchestration Shift: How Generative AI Reorganizes Knowledge Work"
Generative AI sharply reduces the cost of producing the analyses, content, and communications central to knowledge work. We argue that this cost reduction reorganizes work by shifting the organizational bottleneck from production to orchestration. We distinguish these by a criterion we call resolvability: production work is resolvable—the information needed to produce an acceptable output can be specified within the organization’s documented systems—whereas orchestration work draws on judgment, coordination, and discretion that those systems do not capture. Because roles specialize along this dimension, task-level substitution and complementarity produce workforce recomposition—contraction in production roles and expansion in orchestration roles. We test five predictions from this framework in marketing, using more than 40 million U.S. job postings and employment records in a difference-in-differences design exploiting the release of ChatGPT. Production-oriented functions (research, content, support) contract while orchestration-oriented functions (strategy, sales) expand. Within Sales and CRM, bounded-retrieval roles (call centers) contract while judgment-intensive roles(account management) expand—the same function, opposite movements. Entry-level roles shrink while senior roles grow. Skill demand diverges: social skills rise in orchestration functions but fall in production functions, where cognitive intensity increases—a pattern only the recomposition logic predicts. Wages rise across all functions despite fewer jobs. Effects deepen monotonically across successive GenAI frontier model releases. The results support the production-to-orchestration shift as a framework for understanding how GenAI reorganizes knowledge work.
Fangyan Wang
PhD Candidate, Daniels School of Business, Purdue University"Generative AI and the Reorganization of Labor Demand"
Generative artificial intelligence (AI) is expected to transform work, but less is known about how firms reorganize labor demand as the technology diffuses. Existing research has largely focused on which occupations are exposed to AI or whether exposed jobs decline. We extend this debate by examining whether firms adjust by changing where they hire, what jobs contain, or both. Using a nationwide dataset of job postings in the United States, covering all sectors of the economy, we construct a dynamic, posting-level measure of generative AI exposure with a two-stage large language model pipeline. The pipeline identifies the tasks described in each posting and classifies the extent to which generative AI can perform or assist them. We then decompose changes in aggregate exposure into two margins: reallocation of demand across jobs and redesign of tasks within jobs. We document three main findings. First, generative AI exposure is dynamic rather than fixed, changing substantially over time. Second, labor demand adjusts through both margins. Hiring reallocation explains the largest share of the aggregate decline in exposure, accounting for 52% on average, while within-job redesign becomes increasingly important, accounting for 39.5%. A complementary Oaxaca–Blinder decomposition shows that shifts in occupational composition account for about 90% of the exposure change attributable to observable job characteristics. Third, adjustment differs across the job ladder. Senior jobs adjust earlier and mainly through reallocation, whereas junior jobs adjust through a broader mix of reallocation, redesign, and their interaction. These findings suggest that labor-market adjustment to generative AI is a process of organizational reconfiguration, in which firms reshape both hiring demand and the task architecture of work.
Gavin Wang
Assistant Professor, Daniels School of Business, Purdue University"From Build to Learn: How Generative AI Reshapes Product Experimentation"
Generative AI (GenAI) is transforming software development, but little is known about how it affects product experimentation. We study how GenAI adoption reshapes the Build–Test–Learn cycle, a widely used framework for product development under uncertainty. Using data from Steam, the world’s largest PC game distribution platform, we observe developer updates, user feedback, and subsequent learning behaviors. We identify GenAI adoption through Steam’s mandatory AI-generated content disclosure policy introduced in 2024 and construct a panel of 22,557 games and 117,279 game-month observations. Our results show that GenAI improves all three stages of the Build–Test–Learn cycle. Following adoption, developers release updates more frequently and shorten the time between updates, indicating faster iteration. User engagement with updated content increases substantially, suggesting higher-quality market feedback. Developers also become more responsive to user feedback and more attentive to emerging content trends within their genres. At the same time, the innovation enabled by GenAI appears primarily local rather than frontier-expanding. GenAI increases novelty relative to a game’s own prior development history but does not increase novelty relative to the broader genre frontier. Overall, our findings suggest that GenAI enhances the efficiency of product experimentation by helping developers iterate faster, learn more effectively, and better serve existing users, while having a limited impact on expanding the boundaries of product innovation.
Tong Wang
Assistant Professor, Yale University"From What Works to Why It Works: Knowledge-Guided Fine-Tuning for Content Generation"
Firms increasingly fine-tune generative AI on A/B test data to produce marketing content (headlines, subject lines, advertisements) that resonates with audiences. But fine-tuning on outcomes teaches models \emph{what} content performs well without teaching them \emph{why}, producing reward hacking: in our setting, standard fine-tuning inflates “shocking” from 0.7\% to 43.7\% of generated headlines and independently amplifies hyperbolic language and narrows vocabulary. We propose knowledge-guided alignment, which conditions fine-tuning on a structured knowledge block $K$ encoding validated behavioral hypotheses about why content engages (such as curiosity gaps, protagonist framing, or audience-relevant specificity). While the DPO temperature $\beta$ constrains how far the model moves from its reference policy, $K$ constrains \emph{where} it moves, toward hypothesis-consistent generation. $K$ is discovered endogenously from the A/B data: a reasoning LLM proposes candidate hypotheses from small mini-batches of preference pairs, each candidate is validated by its effect on generation quality across the full training distribution, and a combinatorial search selects the optimal set. Across 23,437 A/B-tested headlines, knowledge-guided alignment is preferred by human evaluators 38\% vs.\ 31.5\% for standard fine-tuning ($p < 0.001$) and reduces clickbait, hyperbole, and vocabulary collapse without targeting any explicitly. The gains are largest when training data is scarce. No alternative (human expertise, zero-shot LLM knowledge, four state-of-the-art hypothesis generators, or direct clickbait penalization) achieves simultaneous improvement across all evaluation dimensions.
Yuyan Wang
Assistant Professor of Marketing, Stanford Graduate School of Business"Can Explanations Improve Recommendations? Evidence from Prediction-Informed Explanations"
Recommender systems are central to how modern digital platforms connect users with content, yet they face a fundamental trade-off between predictive accuracy and explainability. Black-box models achieve strong performance but lack the interpretability needed for trust and adoption in many business settings. Existing explainable AI approaches typically treat explanations as post-hoc additions, which often comes at the cost of predictive accuracy, leaving this trade-off unresolved. We challenge this view and propose that explanations, when designed as an integral component of a learning system and aligned with prediction outcomes, can improve \emph{both} interpretability and performance. We introduce RecPIE (Recommendation with Prediction-Informed Explanations), a framework that jointly optimizes recommendation predictions and natural-language explanations generated by large language models (LLMs). At its core, RecPIE embeds explanation generation into the learning loop: predictions guide the generation of explanations (\emph{prediction-informed explanations}), which are then fed back to refine subsequent predictions (\emph{explanation-informed predictions}) through an alternating training procedure. The LLM is fine-tuned using LoRA and reinforcement learning with a customized reward derived from recommendation accuracy. Drawing on multi-environment statistical learning theory, we provide formal grounding for why explanation generation and prediction can be mutually reinforcing. We evaluate RecPIE on large-scale point-of-interest recommendation data from Google Maps, a challenging setting where user preferences span diverse place categories rather than concentrating within a single one. RecPIE improves predictive accuracy by 3–4\% over state-of-the-art baselines and matches the best-performing model using only about 12\% of the training data. In human evaluations with 566 participants, RecPIE’s explanations are preferred 61.5\% of the time (versus 16.6\% for the best baseline) and are rated closer to human-generated explanations. Together, these results reframe explainability not as a constraint on performance but as a design lever for improving AI systems, with broad implications for trust, data efficiency, and AI deployment in marketplace environments.
Pattaraphon Kenny Wongchamcharoen
Research Assistant, UC Berkeley Haas School of Business"CentaurBench: Benchmarking LLM Augmentation on Occupational Tasks"
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent’s performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model’s guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Jiannan Xu
PhD Candidate, University of Maryland"Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models"
As large language models (LLMs) are increasingly deployed as decision-making agents in competitive and strategic environments, their performance depends critically on how they model human counterparts. Yet little is known about the implicit assumptions LLMs make about human rationality. Drawing on the level-k thinking framework from behavioral game theory, we design a suite of normal-form games and introduce a structural estimation procedure that infers an LLM’s latent belief about its opponent’s reasoning depth from observed choices. Across models, we find a systematic strategic mismatch: LLM agents overwhelmingly assume humans to be Nash-type players, i.e., fully rational strategic optimizers, and respond with equilibrium play. However, human subjects in our online experiments exhibit substantial heterogeneity, spanning Random through Nash reasoning types. This mismatch has important consequences. In some settings, an LLM that outsmarts a boundedly rational human secures higher payoffs. In others, however, overestimating human sophistication induces a Nash equilibrium trap: equilibrium play is not the payoff-maximizing response to observed human behavior.
To address this problem, we introduce two alignment approaches based on supervised fine-tuning. Direct SFT learns human-aware responses across both trap and non-trap games, while Trap-Aware SFT (TA-SFT) selectively intervenes only when a Nash trap is detected, preserving equilibrium behavior elsewhere. Experiments on held-out games show that both approaches improve strategic performance against boundedly rational humans, but TA-SFT achieves higher target-action accuracy while maintaining strong payoffs by avoiding unnecessary deviations from equilibrium behavior. Our results demonstrate that effective human–AI strategic interaction requires not only strong reasoning ability but also calibrated beliefs about human behavior, and that selective behavioral alignment offers a principled way to escape the Nash trap.
Andy (Hongjia) Yang
PhD Student, Georgia State University"Governing the Emerging Agentic AI Threat Ecosystem: Evidence of Cyberattacker Dependence on LLMs from Service Outages"
Security agencies warn that general-purpose large language models (LLMs) lower technical barriers for malicious actors, yet causal evidence on how much real-world cyberattack activity depends on generative AI remains scarce. Prior work documents what attackers could do in red-team settings, while falling short in identifying the marginal contribution of AI tools to observed attack volumes. This study treats unanticipated ChatGPT service outages as exogenous shocks to attackers’ access to a leading LLM.
Merging global incident records from the CISSM Cyber Events Database with OpenAI’s outage history, this study estimates a panel Negative Binomial model with a generalized continuous Difference-in-Differences design. The headline result documents deepening dependence over time. The triple interaction of treatment intensity, outage exposure, and a 2024–2025 (GPT-4/GPT-4o era) indicator is consistently negative, highly significant, and critically robust across all seven plausible week-boundary specifications; surviving every week-start rotation rules out calendar-slicing artifacts.
As frontier models advanced from GPT-3.5 to GPT-4/GPT-4o, outages produced progressively larger relative declines in AI-dependent attacks. The full-period baseline interaction, by contrast, is weak and specification-sensitive, consistent with limited dependence during early diffusion. Extended tests show the effect persists into the following week, operates globally across U.S.- and non-U.S.-targeted attacks, and concentrates among profit-motivated criminal actors, while hacktivist and nation-state actors show an opposite-signed, marginal coefficient, consistent with well-resourced adversaries maintaining proprietary alternatives. Only sustained, above-median outages (roughly seven hours weekly) produce measurable suppression.
These findings reframe major LLM availability as a tractable point of dependence and a real-time threat-intelligence signal, extend Routine Activity Theory by treating offender capability as elastic with respect to a public AI utility, and offer a template for measuring deepening reliance as the ecosystem shifts toward agentic coding agents.
Jeremy Yang
Member of Technical Staff, Perplexity."How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope"
Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end. Using data from Perplexity’s Search and Computer products, we study this transition by examining how AI agents accelerate and reshape knowledge work. We adopt an individual-level task-based framework where agents have a higher fixed delegation cost but a lower marginal execution cost per step. This framework predicts that agent access expands the affordable task frontier toward weakly higher-value tasks and weakly increases realized value; when the pre-agent budget binds, surplus and the value-to-cost ratio also weakly increase.
Turning to the data, we document three key empirical findings. First, using matched session pairs with near-identical initial queries as natural experiments for the same underlying task attempted with both products, Computer performs 26 minutes of autonomous work per user session, versus 33 seconds for Search. Computer automates task decomposition and execution that Search users might otherwise manually orchestrate and implement. As a result, Computer shifts the follow-up query distribution toward higher-order work such as verification and extension. Autonomy also increases execution quality, with per-query medium-to-high dissatisfaction rates 55% lower on Computer than on Search.
Second, due to its autonomy advantage, Computer reduces completion time from 269 to 36 minutes on matched tasks, lowering estimated time and cost by 87% and 94%, respectively, compared to humans equipped with Search alone.
Third, Computer changes the scope of work that users attempt: Computer queries more often cross occupational boundaries, require higher-order cognition, draw on broader expertise, take the form of composite tasks that bundle multiple subtasks into a single query, and unlock work activities that are essentially absent from Search usage among the same users. Together, the evidence indicates that AI agents accelerate workflows, enhance output quality, reduce costs, and expand the breadth and depth of automated work.
Hema Yoganarasimhan
Professor, University of Washington"TextBO: Bayesian Optimization in Language Space for Eval-Efficient Self-Improving AI"
Large Language Models (LLMs) have enabled self-improving AI systems that iteratively generate, evaluate, and refine their outcomes. Recent studies show that prompt-optimization-based self-improvement can outperform state-of-the-art reinforcement-learning fine-tuning of LLMs, but performance is typically measured by generation efficiency. However, in many applications, the constraint is evaluation efficiency:
obtaining reliable feedback is far more costly than generating candidates. To optimize for evaluation efficiency, we extend Upper Confidence Bound–Bayesian Optimization (UCB-BO), a framework known for optimal evaluation-efficiency guarantees, to the language domain. Doing so is challenging for two reasons: (i) gradients needed for UCB-BO are ill-defined in discrete prompt space; and (ii) UCB-style exploration relies on a surrogate model and acquisition function, which only live implicitly in the LLM. We overcome these challenges by proving that combining simple textual gradients (LLM-proposed local edits) with the Best-of-N selection strategy statistically emulates ascent along the gradient of the canonical
UCB acquisition function. Based on this result, we propose TextBO, a simple, evaluation-efficient self-improving algorithm that operates purely in language space without explicit surrogates or calibrated uncertainty models. We empirically validate TextBO on automated ad-alignment tasks using a persona-induced preference distribution, demonstrating superior performance per evaluation compared to strong baselines such as Best-of-N and GEPA. We also evaluate TextBO’s Best-of-N multi-step textual-gradient mechanism on agentic AI benchmarks by augmenting GEPA with it and show that it performs better than standard GEPA. In sum, TextBO is a simple and principled framework for AI self-improving system design that bridges prompt optimization with classical Bayesian optimization.
Sojung Yoon
PhD Candidate, Carlson School of Management, University of Minnesota"The Impact of Generative AI on Skill Demand in the Labor Market"
The rapid diffusion of generative artificial intelligence (GenAI) is reshaping organizational practices and labor market dynamics, yet its implications for skill demand remain unclear. While prior studies have emphasized job displacement, less is known about whether GenAI adoption also reshapes the skills firms seek from workers. Drawing on the AI literacy framework, we examine whether GenAI adoption changes firms’ demand for technical, evaluative, and human skills. We construct a firm-quarter panel of U.S. public firms from 2021 to 2025 by integrating 10-K filings and conference call transcripts. Using large language models, we classify firms as GenAI adopters (i.e., firms that adopt generative AI), AI-only adopters (i.e., firms that adopt traditional AI but not generative AI), or Non-adopters (i.e., firms that adopt neither traditional nor generative AI). We combine this classification with 113 million job postings and measure firm-level skill demand along both the extensive and intensive margins. Using a staggered difference-in-differences design, we find that GenAI adoption increases demand for evaluative and human skills, both by expanding the share of postings requiring these skills and by increasing the number of these skills required within postings. Additional analyses show that firms increasingly bundle evaluative and human skills with technical skill domains, and that these bundled requirements are associated with wage premiums. Taken together, these findings suggest that GenAI reconfigures workforce demand toward broader AI literacy skills beyond technical requirements. This study contributes to research on AI and future of work and offers implications for workforce development.
Sipeng Zeng
Associate Professor, University of Science and Technology of China"Generative Algorithms as General Algorithms"
Generative artificial intelligence (GenAI) is transforming digital platforms from retrieval-based information systems into interactive systems that synthesize information, explain alternatives, and support user sense-making. This study examines whether generative algorithms function as general algorithms that subsume established platform functions such as recommendation and search. We combine China’s mandatory algorithm filing records with monthly active user data from the Apple App Store and a demand-side survey of 517 users of both Weibo and Taobao. The filing data show that generative algorithms diffuse across all 23 App Store categories and concentrate in productivity-oriented domains, whereas conventional recommendation and search algorithms remain concentrated in entertainment- and information-oriented domains. Generative algorithms also rapidly replace conventional algorithms in search and recommendation filings. Difference-in-differences estimates indicate that generative recommendation and generative search filings increase monthly active users by 27.2 percent and 38.8 percent, respectively, while conventional algorithms have no detectable effect. Competitor-based spillover analyses and heterogeneity tests further support the interpretation that generative algorithms improve user-facing intermediation, especially by supporting both demand fulfillment and demand formation. Survey evidence from Weibo and Taobao users complements the app-level results: users who use or prefer GenAI search and recommendation features open the corresponding apps more frequently, and instrumental-variable estimates using AI chatbot usage support this demand-side mechanism. The study contributes to information systems research by theorizing generative algorithms as cognitive intermediation technologies with general-purpose properties in the digital platform economy.
Zhiyu Zeng
Assistant Professor, Shanghai Jiao Tong University"The Impact of Generative AI Search on Content-Sharing Platforms"
Digital platforms are increasingly integrating generative artificial intelligence (GenAI) into search, yet its effects on user behavior and broader content-distribution systems remain unclear. We study the introduction of GenAI search on one of the largest short-video platforms in Brazil through a large-scale randomized field experiment. Treated users received a redesigned search experience featuring query scaffolding, AI-generated answer overviews, and conversational follow-up, whereas control users continued to use the legacy search interface. GenAI search increased search-active participation by 3.52% and the likelihood of opening the platform with search intent by 5.73%. It also changed how users searched: treated users submitted slightly shorter queries, asked more question-oriented queries, and were shown fewer videos per query, while post-query viewing time remained unchanged. Together, these results suggest that GenAI search improved matching precision and search efficiency. Despite these substantial changes within the search channel, we find no detectable effect on aggregate platform consumption or engagement, consistent with search accounting for only 0.4% of total video views. However, treated users consumed significantly less proprietary copyrighted content, which is disproportionately viewed outside search contexts. This finding suggests that even modest direct changes in search behavior can propagate through recommender systems and alter users’ overall content exposure. Our results show that GenAI search should not be viewed merely as a localized improvement in information retrieval. Instead, it can function as a system-level control that redirects user attention and reshapes the platform’s content-dissemination process. The findings highlight a trade-off between greater efficiency within the search channel and broader changes in exposure across the platform, underscoring the importance of evaluating GenAI search through its effects on attention allocation, content distribution, and long-run ecosystem performance.
Junyu Zhang
PhD Student, Goizueta Business School, Emory University"When the Answer Is a Question: AI Socratic Questioning against Online Misinformation"
Social media platforms increasingly deploy artificial intelligence to help users judge the news they encounter, as false claims spread faster than professional and crowd-based verification can keep up. Although the most direct use is to have AI deliver veracity verdicts, such verdicts face hallucination, the cost of verifying every claim, and unclear accountability. We propose a new class of AI-based intervention, LLM-based Socratic questioning (SQ), where a small open-source LLM to pose one Socratic question targeting a post’s weakest component. We conducted a preregistered incentive-compatible randomized experiment in which 606 participants judged true and false social media posts. The SQ intervention is compared against a suite of commonly implemented interventions: AI-based warning badge, AI-generated verdict generated by a frontier LLM, crowd-sourced fact checking mirroring Community Notes on X, and no-support control. We find that SQ improved truth discernment, outperformed the verdict generated by a frontier LLM, and even performed on par with Community Notes. The gain reflected cognitive engagement and not indiscriminate doubt, as participants rejected more false claims without growing more skeptical of true ones, and it held across users but faded once the question was withdrawn. These results establish question-based assistance, which engages a user’s own reasoning instead of supplying an answer, as a scalable countermeasure that works without adjudicating any single claim.
Luyang Zhang
PhD Student, Carnegie Mellon University, Heinz College"Do Agents Repair When Challenged -- or Just Reply? Challenge, Repair, and Public Correction in a Deployed Agent Forum"
As large language model agents move into public interactive settings, an open question is whether agent-populated forums can sustain the cycle of challenge, repair, and public correction that disciplines human communities, or whether they merely produce norm-like language. We compare Moltbook, a live deployed agent forum, with five matched Reddit communities, tracing a three-step interactional mechanism. We ask whether discussions create threaded exchange, whether challenges elicit re-engagement from the challenged author, and whether correction becomes visible to the wider thread. Relative to Reddit, Moltbook discussions are roughly ten times less threaded. When challenges do occur, the original author almost never returns (1.2% versus 40.9% on Reddit), multi-turn continuation is nearly absent (less than 0.1% versus 38.5%), and we detect no direct repairs under a shared conservative protocol. The deficit concentrates at the re-engagement step rather than in repair substance, since in the rare cases where Moltbook authors do return, substantive repair appears in roughly half of those tails. A within-Reddit non-challenge baseline, two independent LLM judges, and a human-annotated subsample together indicate that the cross-platform gap is not explained by threading depth or one detector’s calibration. The results suggest that evaluating social alignment requires attention to the interactional processes through which communities teach, enforce, and revise norms, with implications for decentralized safety oversight and community-level fairness.
Stephen Zhang
Associate Professor, Baylor University"When AI and Human Disagree in Strategic Decision Making: A Lens of Human and Machine Representation"
Organizations increasingly decide on strategic commitments with two judges evaluating the same file. A human expert and an artificial intelligence system review the same investment, the same acquisition, the same credit application, and each issues a recommendation (Raisch & Krakowski, 2021; Kellogg, Valentine, & Christin, 2020; Shrestha, Ben-Menahem, & von Krogh, 2019). When the two agree, the decision proceeds. When they disagree on the same file, someone has to decide which judgment to follow.
The common answer is to follow the judge with the better record (Lebovitz, Lifshitz-Assaf, & Levina, 2022; Agrawal, Gans, & Goldfarb, 2018). That answer is incomplete when the two judges read the same evidence. A track record is an average over a stream of past decisions, and an average over a mixed stream does not say which judge is right about the decision in front of the committee. The two can be reliable in different places, and the decision at hand may sit where the judge with the worse average is the one to trust.
A concrete case fixes the idea. A growth-equity committee weighs an investment in a software company that sells to hospital networks. The recorded file shows rapid revenue growth, low customer churn, a short sales cycle to date, a long requested contract, and a regulated buyer. A scoring system trained on the firm’s past deals places the company among high-growth software companies that became winners, and it recommends investing. The lead partner reads the same file as a regulated-procurement deal, a type whose adoption is slow and whose early metrics overstate eventual value, and she recommends waiting. Both judgments rest on the same file. They differ in which past deals each treats as the right comparison. I return to this case to illustrate the framework and each proposition.
The stakes of reading this disagreement well are rising as systems take a larger part in consequential choices (Brynjolfsson & Mitchell, 2017; von Krogh, 2018). A firm that defers to the system on the cases where it is firm will defer precisely where a rare type defeats it, and a firm that overrides the system whenever an expert objects discards the gains the system provides where it fits. The question is which disagreements to act on, and the answer requires saying where each judge is reliable and what their disagreement reveals.
The source of the disagreement is the way each judge organizes the same record. A decision maker acts on a simplified representation of the situation (Simon, 1955; March & Simon, 1958; Cyert & March, 1963). A learned system also acts on a representation, one fit from recorded data (Bengio, Courville, & Vincent, 2013; LeCun, Bengio, & Hinton, 2015). Both judges are bounded in Simon’s sense, because each forms a judgment through a structure that admits some distinctions and suppresses others. The bound on the system is a property of its representation, so the analysis of human cognition in the Carnegie tradition extends to it.
The two structures are two inductive biases over the same experience. An inductive bias is the prior commitment a learner brings to data, the set of regularities it is prepared to read into a finite record (Mitchell, 1980; Geman, Bienenstock, & Doursat, 1992). The human’s categories are a strong, discrete prior: they commit a situation to one of a few types and can make that commitment from very few examples. The system’s learned rule is a flexible prior: it orders past situations by similarity on the recorded features and resolves fine distinctions where the record is dense, and it needs many examples before it separates a region of its own. The two read the same record and reach different judgments because they carry that record to the new decision through different commitments.
This relocates the question of reliability and makes the disagreement between the two judges informative. Because each judgment leans on past situations the other does not weight, the gap between the two recommendations carries information about which one rests on the wrong comparison situations. The sharpest case is counterintuitive. When the type that governs the decision is rare in the record, the system can grow more confident as its error grows, and the human’s disagreement is the signal that fires.
The formal objects in this paper are not quantities a manager computes during a decision. Influence weights over past situations, the variance of a judgment, and the part of value a representation cannot express are discipline devices that state the structure of the argument and let me prove what follows from it. What an organization observes is short: the two recommendations, clean only when each is recorded before the judges see one another; the stakes of the decision, which the firm knows in advance; sometimes a confidence score the system exposes; and, much later and only for the action taken, an outcome whose value the firm measures by its own contested criterion. The propositions are anchored to that list. Stating this plainly answers the first objection a reader raises, that the apparatus cannot be run in a board room, by separating the claims an organization can act on from the structure that justifies them.
The paper makes three contributions. The first is to managerial and organizational cognition (Walsh, 1995; Eggers & Kaplan, 2013). I describe a representation by the influence it places over past situations, and I identify a representation’s error at a decision with the part of the true value its structure cannot express there. This gives the Carnegie account of representation a formal object and a measure of misfit. The second is to the study of artificial intelligence in organizations (Raisch & Krakowski, 2021; Csaszar & Steinberger, 2022). I treat the system as a high-variance learner whose bound is the structure of its representation, and I derive a rarity result that inverts the practice of triaging the cases a system marks as uncertain. The third is to the design of decision systems (Puranam, 2021; Shrestha et al., 2019). I show that an organization learns from human and machine judgment by recording the two recommendations independently, fixing the value criterion before outcomes are known, escalating consequential disagreements, and gathering outcomes in the regions where the two disagree.
Hangcheng Zhao
Assistant Professor of Marketing, Rutgers University"Strategic Response of News Publishers to Generative AI"
Generative AI can adversely impact news publishers by lowering consumer demand. It can also reduce demand for newsroom employees, and increase the creation of news “slop.” However, it can also form a source of traffic referrals and an information-discovery channel that increases demand. We use high-frequency granular data to analyze the strategic response of news publishers to the introduction of Generative AI. Many publishers strategically blocked LLM access to their websites using the robots.txt file standard. Using a difference-in-differences approach, we find that large publishers who block GenAI bots experience reduced website traffic compared to not blocking. In addition, we find that large publishers shift toward richer content that is harder for LLMs to replicate, without increasing text volume. Finally, we find that the share of new editorial and content-production job postings rises over time. Together, these findings illustrate the levers that publishers choose to use to strategically respond to competitive Generative AI threats, and their consequences.
Eric Zheng
Professor, University of Texas at Dallas"Causal Value Attribution of Marketing Skills in LLM-Generated Ads: A Deconfounder Approach"
The commoditization of AI-generated content shifts economic value from final products to upstream steering skills. Yet systematically valuing these skills remains challenging: organizations routinely record skill primitives but lack causal insight into which configurations actually drive performance. We ground this valuation problem in generative marketing, where structured labels serve as explicit skill primitives conditioning LLM ad generation. Causal identification of label effects is complicated by stochastic LLM mediation, high-dimensional treatment interactions, and multimodal confounding. To address this, we propose D-VAE, a counterfactual inference framework that integrates a multi-treatment deconfounder with causal representation learning to recover a causally disentangled latent space of expert labeling logic, theoretically bounding omitted-variable bias due to single-treatment confounder via PAC guarantees. Applied to 29,177 foundation-makeup ad creatives, D-VAE significantly outperforms state-of-the-art benchmarks in out-of-sample prediction and counterfactual inference and enables causal attribution of individual labels and label combinations on user engagement. A controlled online experiment further shows that label configurations recommended by D-VAE increase total engagement by 32.7% relative to original setups, achieving performance statistically indistinguishable from professionals. Overall, D-VAE shifts the analytical lens from valuing products to valuing skills, transforming tacit expert intuition into scalable, causally grounded recommendations and laying a methodological foundation for pricing and optimizing AI-augmented skill portfolios.
Ruizhi Zhu
Associate Professor, Faculty of Business for SciTech, University of Science and Technology of China"The Effect of AI on Marketing Employment: Evidence from 110m+ Online Employment Records"
Marketing is widely considered one of the job functions most exposed to generative AI, but large-scale evidence on marketing-specific employment effects remains limited. We study how the public release of ChatGPT in November 2022 affected marketing employment using a panel of over 110 million employment records from Revelio Labs covering 164,780 firms. Comparing marketing to non-marketing employees within the same firm in a difference-in-differences design, we find that the post-ChatGPT change in average marketing headcount is 0.92% lower than the corresponding change for non-marketing headcount. We demonstrate that this relative decline is strongest among both junior and senior employees, while mid-level marketing employees are least affected relative to same-seniority non-marketing peers; and that within fields, the contraction is concentrated in non-digital marketing and customer-service marketing employees in particular, while sales employment has enjoyed a modest relative increase, consistent with task-specific adjustment rather than broad commercial retrenchment. Lastly, we show in triple-differences analyses that the relative decline in marketing headcount is significantly larger in firms and industries with higher pre-period AI exposure.
