Skip to Content

2026 Financial Markets Conference – Research Spotlight 2 Transcript – May 19, 2026

Research Spotlight 2: The Economic Value of Generative AI

Transcript

Sudheer Chava: Okay, let's get started. I'm Sudheer Chava. I'm a regents professor at Georgia Tech, and today we have a session on the economic value of generative AI. Again, AI has been not only a buzzword but also a large amount of research, and it's in the public for a long time.

So, Greg Schubert from UCLA is going to talk about his views on the value of generative AI, and then I'm going to moderate some questions. So, if you have any questions, please put them in the conference app and then we can ask the questions after Greg finishes his presentation.

Thank you. Greg?

Gregor Schubert: Thank you so much. So, today I will summarize some of my research on the economic value of generative AI. The high-level way, I think, of trying to conceptualize the economic impact of AI is to try to keep three things very distinct, which is that there's, on the one hand, theoretical exposure to generative AI—so, where the technology could theoretically be useful and have an impact. There's actual adoption—are people actually using it? Even if they have some sort of potential, not everyone ends up using it, and it takes a while for people to adopt.

And then there's economic value: How large of a benefit—or a cost—are we actually deriving from this technology? And I'll go through each of these stages and how we can measure these, both for firms and for households.

On the firm side, I'll mostly be drawing on research that's joint with my colleague Andrea Eisfeldt at UCLA and Miao Ben Zhang at USC, and the key idea to measure firm and worker exposure to generative AI is to think about jobs as bundles of tasks. And so, we can take tasks that are within a given job, and we can look at how those tasks are exposed to generative AI. O*NET provides a detailed list of tasks that are contained within each occupation, and so the way we approach this in research is that we score each of these individual tasks against a rubric of generative AI capabilities, of whether or not this particular task is something that generative AI could be useful for and that it could make 50 percent more productive, in theory.

And then, with that in hand you can aggregate that up into occupation-level exposures based on the kinds of tasks that are contained within each occupation, and into firm-level exposures based on the occupational employment structure within each firm, to obtain a measure that basically reflects the average expected share of tasks within a firm that are susceptible to generative AI to some degree—where the technology could provide some productivity benefit. And if you do that—and we did this in 2023, with the generation of generative AI back then—we get a list that looks something like this, where you get companies like IBM at the top, with the highest exposure to generative AI, and then a number of other companies that are all big, white collar service companies in data processing, publishing, telecommunications.

If you do this for occupations, you find that there is a strong positive correlation between wages and generative AI exposure. This might be a bit small to read from the back, but the highest exposure occupations here are computer/mathematical occupations, legal, management, business and financial operations, with office administrative support being one outlier that is relatively low wage, but relatively high exposure.

And an important contrast to previous waves of automation is that here, higher cognitive labor jobs are much more exposed than manual skilled jobs were in past waves of automation. And one important thing to keep in mind is that when I'm saying exposure—and this is oftentimes distorted in public discussions—exposure does not mean that it is necessarily bad for the workers. Exposure here just means that there's some potential use of the technology, and that could be good or bad.

Of course, it is possible that generative AI is used instead of a worker and leads to their replacement, and lower labor demand and wages. But it could also mean that it is a complement, that the technology ends up being used by the worker and makes them more valuable, and perhaps leads to higher labor demand and wages for that particular position.

A lot of this depends on actually things that are happening within the job, which particular tasks are impacted; and so, one thing we like to distinguish in this research is the concept of core tasks versus supplemental tasks. Core tasks are the thing that you do on a daily basis that is actually essential for your employer. That is, the main value-add of you as a human in terms of getting your job done. These might be essential workflows, key decisions, things where you're applying human judgment—and if generative AI were actually able to do those things for you, there would be a higher risk of you being replaced.

On the other hand, there are also many supplemental tasks—things like taking notes in meetings, documenting, transcripts, summarizing, helping with formatting of documents. All of these things are things that we do, but that are not the things that make us essential to our employers; and as a result, if AI can help with those, it might make us more productive and allow us to refocus our time and efforts on the tasks that are actually important—and as a result, might actually increase your value to your employer.

And so, one important thing to keep in mind is that exposure doesn't mean it's negative for the workers, and the replacement of a task does not mean that your entire job gets replaced. It really matters which particular tasks are impacted.

And you can do back of the envelope estimations—and this is all going to be very rough, and ultimately wrong in the long run—but if you just want a rough estimate of how important this all is, this potential labor-side productivity improvement, you might do a very simple calculation like take the total wage bill for an occupation, take this 50 percent productivity improvement number that we use to come up with our exposure measures, and multiply that by the percent of tasks in the job that are exposed, and you'll get very large numbers. If you do this back-of-the-envelope estimate, you'll get numbers in the range of like $1 trillion of potential annual labor value that you could get just from applying 2023 generations of generative AI.

And of course, as these models get better that's only going to be a bigger number. And so, whatever the exact number is, there's very large potential productivity improvements that you would get if you follow these exposure methodologies.

So, do companies actually end up using generative AI? This is also the bottom-up analysis of the potential. What do they actually end up doing? The first thing that we do in our research is to measure how much financial markets think companies are going to benefit from generative AI.

And what we do is we sort companies into portfolios, based on their generative AI exposure, and then form a long-short portfolio that goes along with the companies that have a lot of artificial intelligence exposure, and short the companies that do not. We call this the "artificial minus human" portfolio. And what we find is that when ChatGPT is released, the companies that you would theoretically predict to have a higher exposure to the technology see 0.4 percent higher returns per day during the two weeks of trading after the ChatGPT release.

So, it looks like markets are revaluing these companies, based on the fact that they could use these technologies profitably. And one important note is that you might expect future re-evaluations as other generations of the technologies are released. You can see this a little bit in the graph here.

So, the thick blue line is this long-short portfolio, and you can see it doesn't go anywhere before ChatGPT is released. It jumps up when ChatGPT comes out, and then when GPT-4 is released later in the year in 2023, you get another bump up in terms of valuation for these companies. Of course, at some point this becomes priced in and anticipated, but you can see these companies being revalued based on their technology exposure.

One important wrinkle is that this is not true across the board; this is a heterogeneous impact across companies, where companies that have more preexisting data assets—so, either access to large proprietary data sets, or data capabilities among their workers—are much more likely to see these increases in their stock prices when these new technologies are released. So, markets seem to be imputing that some companies will be much better at actually putting these technologies into practice.

Now, that's what financial markets are predicting. Do companies actually end up using these technologies?

We can use job postings to get one proxy for whether or not companies are actually adopting these technologies, and the way this works is that—I show an example here on the slide for a product specialist associate at a real estate fund—job postings will mention what you're actually doing in a job. And so, at the bottom, for instance, of this job posting, it mentions that this product specialist associate utilizes generative AI—specifically JLL GPT, so they seem to have trained some sort of internal chatbot—to support and optimize specific tasks. And so, you can tell from the job posting, this company is obviously using generative AI, and in particular they're having this particular worker interact with generative AI to some degree.

And so, you can use this to then see if companies actually end up using the technology. And what we find is that there's a strong correlation between this predicted exposure—Do you have potential from the technology?—and companies actually using it.

So, this graph shows companies sorted by their exposure, and then the actual share of their job postings by the end of 2024 that ends up mentioning generative AI-related skills. So, you can see it's a bit nonlinear, but companies that have much higher exposure are much more likely to actually be hiring for generative AI skills. There's definitely some translation here from these theoretical estimates of where the technology is useful, to actually finding companies using those technologies in reality.

However, that translation is not perfect. Some companies somehow find it easier than others to go from "I could use this technology" to actually using it, and so what we find is that this translation from exposure to adoption varies a lot with preexisting workforce technology skills. So, companies that already have a lot of workers with technological capabilities, that have a greater share of managers, and that have generally more complex positions—so, positions that require advanced degrees, and greater experience—are much more likely to translate a given level of potential from the technology into actually using it.

And, again, this aligns with this financial market prediction that I showed you earlier, that companies that have more existing data capabilities are predicted to realize more of the value. So, we actually see some of this in reality. Of course, what this means is that you might get some path dependence, where companies that already have preexisting technology skills and are already leaders in their industry, are better able to realize value from this new technology and might pull further ahead and expand their dominance in their particular sectors.

Why is it so hard for some companies to adopt these technologies? I did a little study where I looked at how companies talk in their earnings calls about investments in generative AI, and what they're saying that they're actually investing in when they're investing in generative AI. What is it that you actually need to do in order to put this technology into practice?

And what I find is that the things that companies most frequently mention when they're talking about what they actually need to do in order to get this technology up and running is, they need to hire some AI vendors and consulting from the outside. They need to invest a lot in compute infrastructure, so they're buying GPUs and hardware; they're building their own custom, proprietary AI models; and they're putting a lot of effort into data preparation, building out data pipelines.

And when you see that, it makes sense of the data that you do need some existing technology capabilities to actually make this work. You can't just sort of off the shelf, take AI, and it does the job for you. You actually need to build out a lot of technological infrastructure in order to realize the value from this.

What does this mean for the workers? We find that generative AI-exposed firms end up reducing the hiring for the most exposed roles. However, this doesn't mean that they reduce hiring overall; they might increase the hiring for new roles that didn't exist beforehand. But we actually do find that this core task distinction makes a big difference, where workers whose tasks are exposed to generative AI see the hiring go down a lot, whereas workers where only the supplemental tasks are impacted don't see a big decrease in hiring for their position. So, it really matters what exactly is going on with what kinds of tasks generative AI is replacing.

And I wanted to put in a little caveat here, that this is a very active literature on the labor market impacts, where people are mostly finding relatively small employment effects so far. And that might be due to the fact that people are restructuring jobs and still experimenting with where generative AI fits in their organization, but it might also be due to the fact that aggregate employment numbers tend to mask this heterogeneity about some jobs being positively impacted and some jobs being negatively impacted.

Okay. That was it for firms. Now let's talk about households. So, I think one important dimension that is oftentimes missed in the discussion is that there's not just a market-side impact of generative AI, but also generative AI use outside of labor markets. All of us are probably aware that you can use generative AI very productively in our personal lives, and that might have a big economic impact that just isn't valued in GDP or in labor market impact. So, this is joint work with Michael Blank at Stanford and with Miao Ben Zheng at USC.

What we do in this work is we take a detailed panel of household browsing behavior—so, we have a panel from Comscore that shows us second-by-second Internet browsing. So, we see what website a machine clicks on, how long they spent on that URL, and then what the next website is that they click on.

And what's nice about that is that it allows us to actually track when a household first starts using ChatGPT. So, we see them going to ChatGPT.com or OpenAI.com, and so we see when they're first starting to interact with generative AI models, and then we can relate that to both their household characteristics and the other browsing behavior that they show, and how they're changing their interaction with the digital economy as a result of now using generative AI.

And what we do is actually we follow a methodology that's somewhat similar to what I just described at the firm level, in that we first score households' exposure to the technology. What households will we expect to use generative AI? And the way we do this is that we label each website with its key activities—What would you actually be doing on this website?—and then for each of those activities we again put together an exposure score by comparing these activities to a rubric of what generative AI can be used for, and determining which of the activities on this website generative AI could potentially substitute. Which websites could you use a chatbot for instead?

And so, once we have this exposure for activities and then for the websites, we can aggregate that into household exposure before generative AI is released—based on your browsing patterns, do you look like someone who could use generative AI productively? And we find that about 11 percent of all browsing duration is on websites that are highly exposed to generative AI, so where you could almost off the bat use a chatbot instead of going to this website—and that exposure tends to be higher for younger households.

And what's really very nice about this is that this kind of theoretical ex-ante exposure, again, actually predicts adoption. So, this time series shows actual usage of ChatGPT in our browsing data over time, and the two different lines—the top line in red shows adoption by households that we would predict to have high exposure, so where their preexisting browsing patterns predict that they can make use of ChatGPT, and the blue line shows it for low exposure households. And what you can see is that there's a big gap, just based on which households actually have potential uses from generative AI, in predicting whether or not they actually end up using it.

So, exposure and adoption, again, relate to one another. So, what are the benefits now of generative AI? What does this mean for the households? So, in order to try to think about household welfare impacts, we first need to think about the welfare impacts of different websites; and so to do this, we categorize websites into productive uses and true leisure uses. And this follows the time-use literature where productive activities we think of as like education, childcare, non-market work, planning travel, shopping, health care research—all those things—whereas true leisure activities online are things like gaming, social media, TV, streaming videos.

And with our detailed browsing data, we can then actually check whether or not you're using ChatGPT in the context of different activities. So, we see the websites you go to right before you use ChatGPT and right after, and so we can make an inference about what kind of task you're currently doing when you're accessing a chatbot. And so, we compare ChatGPT users to a matched sample of demographically similar non-users to see how their browsing patterns vary around these ChatGPT visits.

And what we find is that households use ChatGPT mainly with productive tasks. So, when people are using ChatGPT, they predominantly use it in the context of these things like job search, health research, shopping, or trip planning, and they're much less likely to use ChatGPT around gaming, social media, Netflix—things like that. And that probably makes intuitive sense to all of you, in terms of what these chatbots can help with. And so, ChatGPT is used with productive tasks.

However, when we look at the overall browsing activity of those households when they start using generative AI, we actually find that once they start using generative AI, they end up devoting a lot less of their time to productive browsing, and a lot more of their time to leisure browsing. So somehow, they're using ChatGPT around all their productive tasks, but as a result of that they are then able to spend more time on actual leisure online—on gaming, social media, streaming, and so on.

And what does that mean? How do we interpret that? So, one way to think about that is that households are trying to allocate their time, and there's a general behavioral fact that if you have more time available online, you spend some of that on your leisure activities, some of that on your productive activities—but productive activities tend to behave more like what we would call time necessities. So, if you have twice as much time on the Internet, you're not going to plan twice as many trips, or do your taxes twice; so, you're going to less-than-proportionally allocate some of that extra time to productive activities, and you're going to spend more of it with the fun things you can do on the Internet.

And so, when we observe that leisure browsing rises sharply after you use generative AI, we would expect some smaller (but still some) increase in productive browsing as well, because the increase in leisure browsing suggests to us that your budget constraint has been relaxed. You somehow find yourself with more time you can spend on the Internet.

What we actually observe is that there's barely any change in the absolute productive browsing time on the Internet. And so, the fact that you suddenly have way more time for leisure browsing, and you don't spend any more time on productive browsing—and productive browsing is exactly the thing that you're doing around ChatGPT—means that what must have happened is that you got more efficient at getting your chores done on the Internet, and as a result you were able to reallocate that time into leisure browsing.

And we try to put together some calibrated estimates of this that are very sensitive to various assumptions, but whatever you do you come up with very large estimates of how much more efficient productive online time must have gotten as a result of being able to use generative AI. We find estimates that are plausible in the range of 75 to 175 percent efficiency gains for these productive browsing activities.

And an important thing to keep in mind is that these are very large numbers, relative to what people are finding on the labor market side; and so, we should expect that there are large productivity gains from ChatGPT that are simply not captured by the labor market statistics, but that are happening inside households. And one reason this matters is that we're seeing an emerging generative AI divide with regard to usage of ChatGPT.

These graphs show, on the left-hand side, that there's a big gap in generative AI adoption over time between high-income households at the top, and low-income households at the bottom; and the right-hand graph shows a similar gap for young users at the top, and old users at the bottom. And as you can see, it doesn't look like those gaps are closing; if anything, those gaps are getting wider over time.

And so, if there are these large productivity benefits from generative AI, these big demographic gaps in actual usage should worry us with regard to where these benefits are actually accruing. But as far as I can tell, there isn't much research yet on why we have these large gaps in adoption.

So, to bring this together: I talked about exposure, how it translates into adoption, and how that translates into value for firms and households, and to conclude I just want to note that it's important to keep in mind that there are generative AI impacts in both the market and the non-market economy, where the first one receives a lot more media attention. We can use financial markets to price some of these benefits, but in order to capture some of these household productivity benefits that are not contained in GDP, we'll have to do more elaborate measurement exercises.

The labor market impacts are still uncertain, but one thing that I want to highlight is that there are these large gaps in adoption between firms and between different households. And so, this heterogeneity in adoption means that it's important to track adoption dynamics, not just in the aggregate but by different groups, in order to inform policy—and that we're currently only seeing the productivity impacts from the leaders in technology adoption, not from the laggards. And so, once the laggards catch up, we might actually see even larger productivity benefits.

And then, as a closing thought: It's important that everything I've just shown you is really a moving target. These models get better every month, and as a result the potential productivity impacts are going to be larger and larger over time. And so, any estimates you're producing are always almost immediately out of date, because we're going to get better and better models over time. And so, it's important to produce timely measures of adoption and exposure in order to keep track of where we should see economic value being generated. Thank you.

Chava: Please ask your questions in the conference app. Thanks, Gregor. So, again, he has done a very interesting job with both the exposure, adoption, and where the value is coming in. And I just want to put in, I don't want to go into the details of the papers—one is forthcoming (Journal of Finance) and the other one also probably will be in a top journal. I just want to put in the broader context, because they're moving targets; anything that we talk about artificial intelligence nowadays is probably going to get outdated with the next version of Claude or OpenAI's ChatGPT, so I just want to put some of this in context.

Again, artificial intelligence—is it a buzzword, or is it something real? If you look at a few years back, it's all crypto; and then, agentic AI is the recent one, and probably you combine both crypto and agentic AI, that's going to be the latest buzzword. But is it real, or is it actually the value of it is increasing, so that it can show up in firms and labor markets?

This is just the number of S&P 500 calls where they mention artificial intelligence, and not surprisingly, it's increasing. It doesn't matter which line of business the firm is, everyone talks about artificial intelligence nowadays because if they don't have an AI strategy, the analysts are not going to be very happy about it. And probably, even if it's orthogonal to their business—again, there's a lot of talk, but at the same time it's not just a buzzword.

If one looks at how well the models are doing—this one is from a paper a few years back which introduced the RAG architecture. You look at going back to 1998, the -100 at the very bottom, what you see is where it was on a variety of tasks, but if you look at now (2023; that's around approximately where it ended), many of these models, they're actually doing better than humans at many of the tasks, if you benchmark them. There might be some bench matching, but still the models are doing much better.

Two things: one, they're doing better than humans; and the second is, in terms of the steepness of it. The last couple of years, significant improvement—especially those of you using Claude to others—from the December release onwards, huge improvement in agentic AI. Before, one couldn't completely rely on the code that it generated, but now it automatically does a lot of things. That goes back to the point Gregor was making about all this being in some ways moving targets, how well the models are doing, and the adoption—all of these are local estimates in some way.

This is another recent study which came in better—they do these benchmarks, and when we look at it, this one basically shows how many of the tasks... on the left-hand side you have the scale humans can do, the time it takes humans to do. And then, how many cases can the models do where 50 percent of the time they're successful?

So, as you can see, this is in a log scale. If you do it linearly, you can actually see explosive, exponential performance gain; but you can see that now even the tasks that humans can get done in 16 hours, the models—including the Mythos, which came in recently—somehow, 50 percent of the time they can actually do the task. Again, it's not completely reliable, 50 percent; and the same thing with 80 percent. You do the same analysis 80 percent of the time, the models are getting better.

Again, looking at this extrapolation on one hand in terms of what are those tasks they're doing, including training a classifier, making the large language models more efficient—all these things are in some ways... instead of humans, the models are doing it, and in 80 percent of the cases some of them are getting it right. Again, if we extrapolate it—Where would it be a year down the line?—I guess, again, going back to what Gregor was saying, it's a moving target; that's where it is. Where would we be a year down the line, or two years down the line, if the same pace of progress goes, and what tasks can be done that, again, has a value in both exposure, adoption, and also where does the value accrue?

Very quickly, this is from Anthropic. This shows both in terms of what the model capabilities are, and what the exposure is. If you look at the red one, it's the adoption; even though with many of the tasks the models are doing a very good job, the actual adoption is a bit low. There might be frictions in the market for why it does it, organizational issues; but the adoption is the red one, and it's not there yet.

That's probably one of the reasons why OpenAI and Anthropic both are putting in the forward deployed engineers to actually go in and implement some of these things with some of these companies. But this is going to pick up; again, going to the point of all these local estimates in some way what would be one year, two years down the line.

At the same time, I'm also cautious in terms of the implementation. Again, at the end of the day, the large language models, they have biases and they also are not going to be—you can't eliminate the hallucination. You can mitigate it through a variety of means, but you can't eliminate it. This is one of the papers that I wrote a few years back, and we asked a simple question: What's the revenue of the company? And we ask both over time, and we also ask across the size distribution of the companies.

Not surprisingly, what you find is for large companies the model gets it right, but also hallucinates more. Partly the reason might be, for example, Nvidia—everyone has their opinion on what Nvidia's earnings are going to be, and the model training data set for these large language models is going to come from Reddit and Twitter and other things. So that's why there's more of hallucination also. And for smaller firms, there might not be as much information in the training data, so it's not going to actually show up.

The same thing, this one probably like another limitation of the models. Again, all this, as I said—moving targets. This was two months back: simple car wash test. Somebody says, "I want to wash my car. I live 50 meters away from the car wash. Should I walk or drive?"

Again, anyone can guess how many cases—even the frontier models do that. Out of the 53 models they tested, as of February 18—again, I repeat: as of February 18, because now it might have changed—but out of 53 models, 42 of the models get it wrong. Even the 11 models which get it right, they get it right for the wrong reason, because in the training data set many times they have "walking is better" or "walking is healthier." They give a long, one-page reason on why actually you need to walk to get your car washed.

So, that's why, again, one has to be slightly careful in terms of that. That brings to me again a couple of things. One, where the exposure, adoption, all these are going to be, again, based on when the analysis that Gregor has done two years back; but now if we do it, it might actually show up a bit differently, because the model's capabilities are changing—but also, who it's actually impacting is changing.

The companies with the data assets were thought to be more positively impacted by this AI, but again, as we saw in February and March, in what they call the SaaS apocalypse, many companies—the same companies—were actually getting negatively hit, because Anthropic and OpenAI, when they release the new tools, again, they're getting into the space of these companies. That's why where it would be a year, two years down the line is going to be an interesting question.

The last thing—again, this is something which I don't have answers, and which I have slight concerns; but in terms of how to think about it, again, many of these, when one looks at the technology over a long period of time, there's going to be a pause to job creation because new jobs are going to be created. You can offload some of the tasks with a calculator, for example; you don't need to know all the calculations. You can offload it; cognitive offloading. But when does it actually lead to cognitive atrophy, and also when does it actually lead to cognitive surrender?

In some ways these are the interesting things. Again, lots of code is being generated—millions of lines of code. Everyone is talking about how probably these companies become much more productive, because they're generating millions of lines of code. Everyone is more productive; code generation.

But comprehension of code is rarely happening. Who has to actually review all the peers? That's going to be a question. So, when somebody doesn't code, are we losing the talent to actually review the code? What's the impact on junior people? That's something which, again, I'm not sure what the answer is, and something that I have some concerns.

And going back, another parallel might be GPS. There's some research which basically shows that as you use GPS more, some of the spatial memory is lost. Is it the same thing if we offload, or if we actually completely rely on AI tools—are we losing some of those things?

In Eric and others' papers, they have a lot of positive evidence on how in the context of customer service people are going to be positively affected in terms of using AI tools, especially the low skilled workers; but one of the things that caught my attention is how people are actually depending more on that, even when the answers from AI are slightly less reliable than their own answers. In some ways, they're using AI more; they're the top workers.

So, as the top workers start using it more, where does the training data for these models come from in the future? Because that's something which actually goes back to the adoption and others.

And lastly, the cognitive surrender; there is a new paper which came in a couple of months back which basically talks about people using AI tools, relying completely on them even when they might be wrong. They do an experiment and they find that the cognitive surrender—in some ways, you use it for the first time, second time, third time, it might be right, and you just basically go on using it. And if the AI makes mistakes—which it will, which I have shown a couple of examples of that—and then people don't actually question critically to that.

So, bottom line: again, Gregor has a couple of interesting papers, and where I would put it all in context is, these are moving targets. They're local effects, in some ways. The models are becoming so good over time; they continue to improve, but at the same time you can't completely eliminate the errors in the models. But the adoption, and how people adopt—both the firms, and the people adoption—is going to change. But also, both there might be positive effects and also negative impacts, some of these things that I've mentioned.

Okay; I welcome questions. Please put the questions in the conference app, and then we can go from there. Thank you.

Okay. So, there are a bunch of questions; maybe if we start with, Kingsley has a question about: Are we making too much of the efficiency gains of AI? Did Google search and Excel spreadsheet not give us the same magnitude of gains? Is it very different from that?

Schubert: That's a good question. So, I remember actually getting a similar question to this, early in 2023, in a seminar.

So, I think there are elements of generative AI that are similar to Google and spreadsheets in the sense that they are very circumscribed tools that can be used for particular things. For instance, if you just thought of generative AI as doing text summarization, and text creation, and maybe some grammar checking and so on; I would probably agree. That's a technology that is highly productive and very useful in white collar work, but it's not going to be transformational.

I think the reason why this technology is probably different from some of these previous technologies is exactly what you were mentioning earlier, which is that, especially since last November, we now have this rise of agentic AI tools, where suddenly the ability to chain together different generative AI tool calls in a way that, when you do ten of them in a row it doesn't fully fall apart. That agentic usage of AI means that suddenly you can throw in high-level questions—or high-level instructions, as in "Go and research all the companies in the S&P 500 and tell me the biggest issue they highlighted in their earnings call last year."

That sounds like an insane task three years ago, and would be several months of work for an analyst, and now you can sort of throw that into an LLM and come back in the morning after it's done the work overnight, and it can provide you with a full report on the answer. And the fact that they can now do these complex, higher-level tasks, structuring their own intermediate steps, I think that is the real productivity gain that goes way beyond being a tool that is only used in a narrow context, but rather now they can take over much more complex chains of tasks.

Chava: So, another question from Larry is about: Your presentation mostly focused on the replacement of existing tasks, but thinking about the Jevons paradox, the increased efficiency will result in increased demand for the related service. So, what are your thoughts, in the context of AI, about this?

Schubert: Yes. So, the Jevons paradox question is the most salient point with regard to what this is going to do to workers. The idea that if you have an increase in efficiency, for instance, of developers, maybe that actually means that demand for code and programming overall goes up so much that we're going to need more developers over time.

My research doesn't take a strong stance on this, and so I'm not sure if I can easily speculate on what's going to happen. I think it's definitely going to be true that in some pockets of the economy, we're going to see demand go up because suddenly, there's more things to be done, and as a result people are going to demand more of the related goods.

So, for instance, one thing you might imagine is that generative AI suddenly makes it much easier for everyone to create customized tools. Everyone can suddenly create customized dashboards for your life; I have a student in my MBA class, for instance, who just vibe coded himself a custom dashboard to keep track of his children's sports games and when he has to be where. And then he added a little prediction engine for whether or not his children are likely to win with their teams in the games they're going to.

And so, you can suddenly create these customized tools, but what that probably means is that there's going to be maybe even more demand for people who are experts in terms of the graphics aspect of that, and giving people advice or consulting on how to create good versions of these. So, you get a much larger demand, and then the market expands—in some sense massively—for some of these customized specialized tools, and maybe that will actually increase demand for some of the related experts.

But more generally, I think this is an empirical question where we need to go one by one, occupation by occupation or industry by industry, to see where for some of them demand's going to go up, for some of them demand's going to go down.

Chava: There's lots of vibe coding. I saw something, again, on LinkedIn where somebody says that my boss vibe coded an app, and now I have to spend 40 hours to actually fix it. So, there's going to be a lot of that.

Andréa is talking about maybe a different aspect of this: Everyone seems very focused on addressing cyber vulnerabilities because of the Mythos that came in (Anthropic). So, will the burst of hiring and spending on cybersecurity temporarily reverse the productivity and labor force implication of generative AI?

Schubert: Maybe. I think, in my mind, this falls generally under this "J curve" idea, of any adoption of a new technology requires you to first invest a lot in restructuring all your existing organizations and processes to align with this new technology, and so in the short run there might actually be a sort of negative impact on observed productivity because you're doing a lot of internal investments into this new intangible capital. The cyber piece probably means that, in the long run, a lot of our cyber defenses are a lot more robust, but in the short run that means a lot of investments in defense capabilities.

I would like to relate that back to the previous question around tasks. I think one thing that becomes very obvious there is that that's going to create a lot of demand for workers related to that; you're going to need people who actually work with generative AI in order to implement these cyber defense capabilities, and build all of those systems.

And so, I like the phrase that generative AI does things middle to middle rather than end to end, because what that means is that you need workers to help, in terms of the data pipelines and actually preparing your organization structures to use generative AI productively. Also, then, you need a lot of workers who do the hard job of validating and checking that what comes out of your generative AI systems is actually useful, and so there are going to be new positions that are being created, on both ends of processes, in some sense. And of course, there's going to be some new processes, like certainly these cyber defense capabilities.

And so, we're going to see a lot of restructuring, but there are definitely going to be new positions that will pop up.

Chava: Hopefully. And this is a question from David about: How do you use AI in the classroom? All of us struggle with that nowadays.

Schubert: I don't know if I have the perfect answer to this, but one thing I do is I actually teach courses related to how to teach managers how to use AI agents well, and so one thing I do in that is that I take it as given that all of my students will be using AI in the classroom, and I have in-classroom exercises and out-of-classroom exercises that actively tell them how to use AI chatbots and agentic tools in a good way.

So, for instance, that encourages them to figure out how to do this validation, rather than just throwing in a question and saying, "do this research task for me," the answer is 42, let me take that away and I'm just going to run with this. How do you actually ask good questions of your chatbot, to figure out: Did it do the analysis correctly? Can you validate intermediate steps? Can you produce some human benchmark to compare the AI classifications against, for instance—have some standards that you can use to assess whether or not chatbots are getting better at the task you're setting them?

And so, there is actually an almost scientific process, in terms of validating the output from these chatbots; and part of what I try to teach my students is to have good, I guess "process hygiene" when interacting with chatbots.

Chava: Yes. In my classroom, I let them use whatever they want, but at the end of the day I could call them and ask them to come and present—because ultimately, the human is responsible, and they get to take the blame. Somebody needs to take that.

Schubert: You're doing reinforcement learning on your students.

Chava: Yes. And maybe this question about, in terms of: What does it mean for, again, a lot of PhDs, a lot of process, and there's a lot of—and, again, at UCLA you just ran an experiment where the AI-based papers were completely written using AI tools. What does it mean for professors, and what does it mean for economic research and finance research? So, what are your thoughts on that?

Schubert: So, Sudheer was referring to the fact that we just had a conference at UCLA called Human x AI Finance, where we actually encourage people to write papers with AI, and then we built a system that reviewed all the papers with AI, and we basically just pushed a button and the system told us which four papers to invite to the conference. The idea was to try to test out this idea that, in the future there's going to be a flood of AI-generated papers, and we need to find ways to filter through that and figure out which papers are actually any good.

And I think it comes back to actually the same idea of AI doing things middle to middle, and then we need to create new tasks on the ends, basically—which is that the role of the researcher is going to move towards preparing data, finding new data, and preparing interesting questions and figuring out what interesting questions are. AI takes over some of the execution parts of research; it can write code, it can clean some of the data, it can do some of the literature review for us, but at the end of the day then researchers need to invest more time at the end—again, in terms of validating the output, actually putting it together, and also communicating it.

And so, the researchers' role of structuring what good ideas are, and then figuring out: Was the analysis done correctly, and what does this mean for the world? It's not that dissimilar to how I think a lot of senior faculty interact with research assistants in the past, and so in some sense the generative AI then enables a lot of researchers to get more work done—but also, they need to be careful about how they do this validation at the end.

Chava: So, were you surprised with any of the papers that came out of that conference?

Schubert: I think I was surprised by the fact that we found that, actually, still there is an experience gradient, to some degree—that people who have more experience working with these tools and are more experienced researchers, ended up producing better AI-assisted papers than people who did not have this experience. So, I don't think the role of experience entirely goes away.

On the review side, using AI as a reviewer, one thing we actually found that was quite surprising was that AI as a reviewer differs in its opinion from human reviewers. So, we had some human reviewers give us a benchmark score on papers, and we checked how our AI review system compared to that. We had multiple human reviewers, and what we actually found was that the AI score was a better predictor of the average human opinion than any of the individual humans—so, variance across our human reviewers, which was much larger than between the AI and the average of the humans.

And so, one thing to keep in mind is that our past systems of human review weren't perfect either, because humans are also highly variable in terms of how we assess things. And so, AI might provide a biased but lower variance way of reviewing papers, for example.

Chava: It's very interesting. There's going to be a flood of papers coming with AI now. So, thank you, Gregor, for your insightful comments on this. Please be seated, because the next session starts immediately after this one. Thank you.

Related