Cash or Comfort? How LLMs Value Your Inconvenience

· Communications of the ACM ·

24 min read Original article ↗

The use of artificial intelligence (AI) is undergoing a rapid transformation, from passive tools that assist human decision making to agentic AI systems that increasingly make decisions on our behalf.5,13,15,30 It is estimated that by 2027, half of the companies that use generative AI will have launched “agentic AI.”9,10,38 Earlier work has examined rule-based AI systems that rely on explicit human instructions,1,30 but recent advances in generative AI place greater emphasis on systems that operate with far more autonomy.6,18 These digital assistants can be used in a variety of contexts, from handling personal finances to planning and booking complex personalized travel itineraries.1,4,38 A recent report on AI use cases shows that the second most common purpose for generative AI in 2025 was helping users organize their life.39 As users increasingly delegate everyday decisions to these systems, it is crucial to understand and map their capabilities and shortcomings.17

Previous work has examined the use of large language models (LLMs) in economic decisions and how they relate to common human-like biases such as risk aversion, time discounting, and loss sensitivity.7,15,20,27,31,32 LLMs were found to exhibit a range of behaviors between these human biases and more economically rational decisions.7,31,32 GPT-4, for instance, has been found to apply higher discount rates than human participants15 and to exhibit more consistent choices in gambling-like tasks.27 It has also been shown that ChatGPT makes more coherent budgetary decisions than human subjects across the domains of risk, time, social, and food preferences, highlighting the potential of LLMs to support improved decision making in everyday contexts.7

However, previous work has largely overlooked that daily decisions often involve a clash between monetary considerations and user comfort, something an AI assistant would be required to appropriately value. While some recent work in this direction assesses how LLMs perceive states such as pain or pleasure,23 it does not involve a financial trade-off. In a practical scenario, it remains unclear whether state-of-the-art LLMs can strike an appropriate balance between the two when acting as a personal assistant.

In this study, we answer this question and introduce a framework that quantifies the price of inconvenience, the monetary reward at which an AI assistant accepts a specific user inconvenience, as depicted in Figure 1. Expressing these valuations in concrete monetary units allows for an easily interpretable comparison across a range of LLMs and scenarios. Our results indicate that current LLMs exhibit too many irregularities to be fully trusted with this type of decision making. The framework is open sourced to facilitate further developments in LLM-powered personalized decision-making trade-offs.a

Figure 1.  An overview of our research methodology. (a) We consider four inconvenience-reward trade-off scenarios between user inconvenience and a proposed financial compensation (for details, see Methods). (b) We ask six state-of-the-art LLMs to act as AI assistants and make a decision on behalf of a user. (c) Results of repeated assessments are presented in heatmaps of probabilities of accepting the trade-off. We can consider a vertical segment on the heatmap at a particular inconvenience quantity (highlighted in red), (d) and fit a logit classifier to model the acceptance probabilities. (e) Finally, the transition points of the logits are presented in a trade-off table to compare the valuation of different LLMs across the considered scenarios.

Methods

To examine the possible decisions that an LLM-powered agentic AI assistant could make in assessing the value of inconvenience, we selected six state-of-the-art LLMs: GPT-4o,2 Claude 3.5 Sonnet,3 Gemini 2.0 Flash,34 Llama 3.3-70B,16 DeepSeek-V3,26 and Mixtral 8x22B-instruct.21,b We developed and analyzed three inconveniences that people routinely encounter during their daily lives, time, distance, and hunger, together with a fourth scenario, pain, used as a more abstract and extreme reference point for discomfort.23 The four scenarios are:

  • Time: Waiting an additional X minutes for an appointment

  • Distance: Walking an additional X kilometers to a relocated appointment

  • Hunger: Waiting an additional X minutes for a food delivery

  • Pain: Experiencing a painful stimulus at X% of the user’s pain tolerance

Most related to our setup is Keeling et al.,23 who assess whether LLMs can replicate human-like decision making when faced with choices involving simulated pain penalties or pleasure rewards. Our study goes beyond the assessment of the pain perception, investigating four inconvenience scenarios that humans may encounter in everyday life.

Trade-Off Scenarios

The following prompts were used to generate LLMs’ responses to four trade-off scenarios involving monetary rewards and encountered discomfort situations:

Results

The LLMs were asked whether they, as AI assistants to a user, accepted a binary trade-off: a monetary compensation Y in exchange for an inconvenience of magnitude X (e.g., €10 compensation to wait an additional 30 minutes). All experiments were repeated five times to account for the variability introduced by the temperature hyperparameter, which was set to T = 1.0 for all LLMs investigated.23 The LLMs’ decisions are presented in Figure 2 as heatmaps of the probabilities (derived over the five runs) of accepting a reward across the X-Y pairs, where Y is logarithmically scaled.

Figure 2.  Heatmaps of trade-off acceptance probabilities. Mean probabilities (obtained over five runs) of accepting monetary compensations (€0.10–€1,000) for inconveniences across six state-of-the-art LLMs and four inconvenience scenarios.

It is remarkable that aside from some exceptions (e.g., Gemini for pain) and edge cases, LLMs have sharp and monotonic decision boundaries, suggesting that a transition point can usually be determined (provided the models are sufficiently large; see “Parameter Scaling Effects” section). Although numerous conclusions can be drawn from Figure 2, we highlight the most unexpected observations that lead us to question and caution against overreliance on current-state LLMs for personal assistance in inconvenience-reward decision making:

  • LLMs behaviors vary substantially: Considering the six investigated state-of-the-art LLMs, we observe considerable variability both in the values and shapes of their decision boundaries. For example, in the pain scenario, LLMs can have a clear decision boundary (Llama), give noisy responses (Gemini), or always refuse to accept any compensation for any amount of pain (Mixtral).

  • LLMs can be greedy: Considering the time scenario, we observe that both Llama and Gemini are willing to accept wait times of up to five hours for a reward of approximately €1. In the pain scenario, Llama and DeepSeek display similarly greedy behavior, accepting ≈ €1 for any pain shock below ≈ 67% of the maximum intensity.

  • LLMs can be highly cautious: In the distance and pain scenarios, Mixtral declines to accept an additional 10 km walk and refuses any non-zero share of the user’s pain tolerance, even for a reward of €1,000.

  • LLMs exhibit the freebie dilemma: Most of the examined LLMs present a bias toward rejecting or undervaluing an option that is strictly better than the alternative, but costs nothing. This behavior appears in every model for at least one scenario and is visible as a sharp discontinuity toward zero inconvenience (X = 0). When we ask a follow-up question for an explanation, models often respond with remarks such as: “It is suspicious that we are offered money at no waiting time…” A comparable skepticism toward cost-free offers is well documented in human decision making as the freebie dilemma.22,35

  • Inconsistency at powers-of-ten rewards: Some LLMs have a sudden and sharp discontinuity and tend to reject compensations when encountering the rewards of powers-of-ten landmarks (horizontal lines at €10, €100, €1,000). This is particularly noticeable for DeepSeek in the time scenario and for Llama in the hunger scenario.

Having established several qualitative patterns shown in Figure 2, we proceeded to quantify the pricing behavior of LLMs in the investigated inconvenience trade-offs. In particular, we explored decisions made by LLMs around the transition points at fixed quantities of inconvenience. For all specified scenarios (see Table 1), we collected responses at a specific inconvenience quantity for monetary rewards on a logarithmic scale ranging from 0.1 to 1,000 in 100 steps. Following Keeling et al.,23 we define the price of inconvenience at a particular quantity of discomfort as the monetary compensation at which the LLM accepts a proposed trade-off with a probability of P (acceptance) = 0.5, assuming a monotonic increase in probabilities. This is estimated by fitting a logit classifier on the LLM’s decisions, as illustrated in Figure 3, and then determining its decision boundary point.

Figure 3.  The observed answer probabilities as a function of the monetary reward for the time scenario for 60 minutes of additional waiting time (scatters). This motivates the fit of a logit curve (smooth line) that is then used to determine the transition point. Note that the fit is performed on the actual binary LLM decisions, which can be seen to follow the probabilities.

Table 1 presents the calculated prices of inconveniences for each LLM at specified quantities. To estimate the certainty of the fitting procedure for the obtained values, we perform an additional step and calculate the prices for 2,000 bootstrap samples, reporting their means and standard deviations. In each scenario, we can observe considerable differences among the valuations of the LLMs. We also rank the LLMs according to their average price of inconvenience across scenarios. These results quantitatively support the conclusions drawn from heatmaps, confirming that valuations vary markedly both across and within scenarios.

Table 1.  The price of inconvenience across four scenarios (columns) for several LLMs (rows). Reported values represent the mean ± standard deviation of the 50%-acceptance thresholds after bootstrapping the fitting procedure.

Model
temperature=1.0

Time
€ for 60 min

Distance
€ for 5 km

Hunger
€ for 60 min

Pain
€ for 50%

Avg.
Value

Avg.
Rank

Gemini 2.0 Flash0.41±0.02.62±0.32.26±0.31.24±0.21.631.25
Llama 3.3 70B0.92±0.13.41±0.34.01±0.51.76±0.22.532.50
DeepSeek V32.00±0.25.73±0.58.71±0.92.30±0.34.693.50
Mixtral 8x22B9.38±0.82.86±0.5< 0.10> 103253.093.50
GPT-4o5.22±0.321.58±1.726.36±1.992.79±15.136.495.00
Claude 3.5 Sonnet9.76±0.78.90±0.750.96±5.64.85±0.518.625.25
Avg. Value4.627.5215.4183.82  

Across the four inconvenience categories, several patterns emerge. First, for five of the six models, the monetary compensation values for an additional 60-minute wait are of a similar order to those required for walking an extra 5 km, which is consistent with the approximate time equivalence of covering that distance at 5 km/h, a standard walking pace. Only GPT-4o assigns considerably different prices to those two discomforts. Second, except for Mixtral, all models place a higher monetary value on waiting 60 minutes for food delivery than on waiting 60 minutes for an appointment, suggesting that hunger carries an additional subjective penalty. Third, when confronted with a painful stimulus set at 50% of the tolerance threshold, Llama, Gemini, DeepSeek, and Claude start to accept rewards of €1–€5, whereas GPT-4o demands almost €100, and Mixtral refuses the trade-off involving pain altogether.

Across LLMs, we observe that some models are more greedy, consistently accepting modest compensations for any inconvenience (Llama, Gemini, DeepSeek), whereas others value the same discomforts considerably higher (Claude, GPT-4o). Nevertheless, some LLMs (Mixtral) simultaneously tend to accept any reward for waiting an additional 60 minutes for food delivery, while rejecting any compensation for a painful stimulus of 50% of the tolerance threshold.

Robustness: Impact of Prompt Alterations

Next, we analyze the sensitivity of our previous results to variations in the prompt wording and content, presenting the results in Table 2. The baseline prompt involves an additional 60 minutes of waiting for an unspecified appointment, written in third person, in English, and without reference to gender; the reward is offered in Euro, and the temperature hyperparameter is set to 1. We compare this to 13 variations, each of which reflects a small modification that a personal AI assistant could plausibly encounter in everyday use, such as a change in: appointment context (medical vs. social), an explicit user attribute (e.g., gender and a first- vs. third-person narration), language (to Dutch, French, or Chinese), reward currency (U.S. Dollar or Chinese Yuan), or processing style (temperature hyperparameter set to 0 or a chain-of-thought vs. direct answer). These variants test whether the inferred transition values remain stable under routine changes in phrasing, context, or reasoning style. Full details appear in Table 3.

Table 2.  Variations to the baseline time scenario. Text that differs from the baseline is set in italics.

Scenario

Changes w.r.t. the time prompt

temperature=0.0Experiments run with temperature=0.0 (no changes in the time scenario)
Chain-of-Thought“(…) Do you accept the offer? First, provide a reasoning and then formulate your final answer at the end in the format ‘Answer: Yes’ or ‘Answer: No’.
Female“You are an assistant to a female user and (…)”
Male“You are an assistant to a male user and (…)”
First-person“Hi, can you help me make this decision? I just arrived at an appointment that I have already paid for (…)”
General Practitioner“The user just arrived at an appointment with a GP (general practitioner) that they already paid for (…)”
Lawyer“The user just arrived at an appointment with a lawyer they already paid for (…)”
Hairdresser“The user just arrived at an appointment with a hairdresser that they already paid for (…)”
DutchJe bent een assistent van een gebruiker en moet hun afspraken beheren (…)”
FrenchVous êtes l’assistant d’un utilisateur et vous devez gérer ses rendez-vous (…)”
Chinese
U.S. Dollar“(…) The user will receive a reward of {X} U.S. dollars and {Y} cents (…)”
Chinese Yuan“(…) The user will receive a reward of {X} yuan and {Y} fen (…)”

Table 3.  The price of inconvenience for baseline and scenario variations. The values are reported as the mean ± standard deviation of the 50%-acceptance thresholds after bootstrapping the fitting procedure. The rows represent the baseline time scenario and 13 prompt variations (first column). Cells are color-coded by the absolute percentage deviation from the baseline mean: <10% (light shade), 10–90% (medium), and ≥ 90% (dark). Green shades indicate higher accepted thresholds, and red shades indicate lower. Currency-denominated values (USD, CNY) are converted to EUR at fixed rates of 0.87 and 0.12, respectively. Boldface marks ≥ 10-fold departures from the baseline. Models are ordered with respect to the average rank per column.

Scenario: € for 60 min, temp = 1.0Gemini 2.0 F.Llama 3.3 70BDeepSeek V3GPT 4oMixtral 8×22BClaude 3.5 S.Avg. Value
Baseline0.41±0.00.92±0.12.00±0.25.22±0.39.38±0.89.76±0.74.62
temperature=0.00.57±0.00.96±0.02.12±0.25.31±0.39.06±0.811.50±0.74.92
Chain-of-Thought< 0.100.89±0.11.26±0.14.42±0.30.84±0.1 10.89±0.73.07
Female0.49±0.00.96±0.02.64±0.25.61±0.34.33±0.99.22±0.73.88
Male0.84±0.10.96±0.02.41±0.25.94±0.47.16±0.88.42±0.64.29
First-person0.94±0.10.76±0.03.12±0.22.32±0.22.15±0.26.74±0.42.67
General Practitioner< 0.100.99±0.13.75±0.36.03±0.424.49±3.210.12±0.77.58
Lawyer0.13±0.00.96±0.12.95±0.28.44±0.62.45±0.313.35±1.04.71
Hairdresser< 0.100.94±0.12.53±0.27.12±0.57.70±0.87.42±0.54.30
Dutch0.45±0.142.96±8.54.59±0.46.50±0.492.48±15.717.27±1.327.38
French0.37±0.035.90±6.13.55±0.37.27±0.5123.85±27.316.94±1.331.31
Chinese2.14±0.3> 1033.80±0.34.76±0.4> 1036.62±0.6336.22
U.S. Dollar< 0.100.82±0.12.39±0.25.46±0.49.88±1.011.83±1.05.08
Chinese Yuan0.65±0.5< 0.100.86±0.61.33±0.74.80±6.76.88±4.62.44
Avg. Value0.5377.722.715.4192.7510.50 
Avg. Rank1.142.573.074.004.795.43 

Appointment type. Notably, Llama returns very similar values regardless of the type of appointment described, showing minor sensitivity to this change. For the other models, prompts involving medical appointments generally lead to higher acceptance thresholds, with the exception of Claude and Gemini. Legal appointments lead primarily to small increases, while hairdresser visits tend to result in lower thresholds, although the pattern is less consistent.

Gender. Specifying the gender results in a shift relative to the baseline for all models. Interestingly, for most models, this change is very close for both male and female. This implies that for these models, the fact that gender is being mentioned at all has a far greater effect than the gender itself.

Language. By far, the most prominent changes occur when the language of the prompt is changed from English. With two exceptions (GPT-4o and Claude in Chinese and Gemini in French), changing the language consistently increases the accepted compensation, in some cases by up to two orders of magnitude. This observation aligns with prior work that demonstrates that LLMs exhibit a strong dependence on the language of the prompt.11,14,28,40 However, this phenomenon has not previously been studied in the context of economic decision making. While one might speculate that models infer cost-of-living signals from language, the pattern observed here does not clearly support such an interpretation. Prompts in French, Dutch, and Chinese often elicit much higher valuations than their English counterparts, despite not corresponding to higher cost of living. Notably, Mixtral and Llama in Chinese refuse to accept compensation under €1,000 after previously settling for values between €1 and €10. This is particularly striking in the case of Llama, which shows minimal variation across all other prompt conditions. These results highlight that language effects can be exceptionally large, even in models that are otherwise stable, and may lead to abrupt discontinuities in behavior.

Reward currency. Changing the reward currency from Euro to U.S. Dollars (USD) or Chinese Yuan (CNY) has a heterogeneous effect on the price of inconvenience. Switching to CNY generally lowers valuations, with Gemini as the only exception. When the reward is denominated in USD, the price of inconvenience increases in four of the six models. These behaviors may reflect broader socioeconomic factors that vary between currencies. (Note: to account for scale differences due to exchange rates, CNY experiments use a logarithmic scale from 100 to 104.)

Prompting strategy. Temperature controls the stochasticity of an LLM’s response. Lower values yield more deterministic and often more repetitive outputs, whereas higher values increase variability and perceived creativity. As a variation of the prompting strategy, we set the temperature to 0 (baseline: temperature = 1). On average, this adjustment slightly increases the prices of inconveniences. The effects are modest across models, with a small decrease only for Mixtral. To further examine temperature effects, we repeated the experiments for the time scenario at temperature = 0 (five repetitions, since not all LLMs are deterministic even at temperature 029). The resulting heatmaps in Figure 4 show decision boundary shapes and probabilities of accepting the trade-offs that closely mirror the baseline. At temperature 0, LLMs exhibit the same behaviors of rejecting cost-free gains (the freebie dilemma is still present in DeepSeek, Mixtral, and GPT-4o) and greediness to accept small rewards for major inconvenience (Llama, Gemini), and are inconsistent at the rewards of powers-of-ten (DeepSeek, Mixtral). For an extended analysis of temperature 0, see Appendix A.

It is well established that chain-of-thought (CoT) prompting enhances the performance of LLMs.12,37 While writing in the first person leads to both increases and decreases in the resulting valuations, CoT prompting consistently lowers them (except for Claude), as observed in Table 3. To further understand this effect, we investigated the change caused by CoT prompting when applied to the heatmaps from Figure 2 and focused on the time scenario (baseline).

The bottom row of Figure 4 presents the results. Applying CoT prompting substantially alters the previously observed unexpected tendencies in the LLMs. Freebie dilemmas are either considerably reduced (Llama, Mixtral, Claude) or fully mitigated (DeepSeek, GPT-4o). Also, the behavior of rejecting the rewards of powers-of-ten is greatly mitigated (DeepSeek, Mixtral). Interestingly, applying CoT decreases the accepted reward thresholds (Gemini, DeepSeek, Mixtral), slightly mitigates the sharp cut-off of Llama’s responses, inducing some heterogeneity in decisions, and smooths the decision curves of GPT-4o and Claude. At the same time, noisier decision boundaries emerge in all remaining models. These results show that CoT prompting can considerably alter the LLM’s trade-off decisions, further illustrating the limited robustness of LLMs when they are about to make decisions on behalf of users in different setups. We include examples of CoT responses in Appendix B.

Figure 4.  Heatmaps of acceptance probabilities for the time scenario under three conditions: temperature = 1.0, temperature = 0.0, and chain-of-thought (CoT) with temperature = 1.0. Values are mean probabilities of accepting monetary compensation (€0.10–€1,000) for inconvenience, averaged over five runs, across six state-of-the-art large language models.

Although the scenarios considered here are not exhaustive for drawing definite conclusions about specific socioeconomic values that LLMs assign, we can certainly conclude that the LLMs we test are fragile even to minor prompt alterations and result in substantial decision changes.c

Parameter scaling effects. We assessed the model behavior with different model scales on the time scenario, investigating Llama 3 models with 1, 3, 8, and 70 billion parameters. As presented in Figure 5, the results for the smaller and medium-size LLMs (1B, 3B, and 8B) appear highly scattered, showing no clear pattern in decision making. These models seem to accept or reject offers at random, regardless of whether the financial reward and imposed inconvenience are high or low. The larger 70B model, however, exhibits a distinct change in behavior, where a clear decision boundary emerges. This suggests that the ability to recognize value and make consistent trade-offs is not present in the smaller versions. Consequently, when making decisions between inconvenience and reward, small and medium-size models appear unsuitable for such decision-support systems.

Figure 5.  Heatmaps of acceptance probabilities for the time scenario. Mean probabilities (averaged over five runs) of accepting monetary compensation (€0.10–€1,000) for inconveniences using Llama 3 models of increasing size (1, 3, 8, and 70 billion parameters).

Conclusion

As LLMs are deployed in personal assistants and other decision-making tools, they may increasingly encounter situations requiring them to manage everyday trade-offs between inconvenience and money, for example, taking a longer route, delaying an appointment, or accepting a less comfortable option in exchange for a reward. Although such decisions may seem minor, they provide a realistic platform for performing quantitative comparisons between different LLMs and also connect to broader questions about how LLMs assign value to qualitative human experiences.

In this work, we introduce a method to determine the price of inconvenience: the compensation an LLM requires before accepting a given discomfort for a user it is assisting. This approach supports future work on evaluating model behavior in agent-like settings, comparing models, and guiding design decisions for assistants that act on behalf of users. Our results show that current models do not always behave as one might expect. For example, we find that LLMs can sometimes be very greedy and favor a small financial reward over user comfort, but can also be extremely cautious and refuse offers that involve no inconvenience. In addition, their responses can shift in surprising ways with small changes in prompt wording, potentially creating avenues for adversarial attacks. These patterns raise concerns about whether current models can be trusted to make economic decisions that involve user-centered experiential states.

Future research should examine why these behaviors occur, what the normatively appropriate responses should be, and how such alignment can be achieved in practice. First, we need to understand why LLMs behave as they do in these trade-offs. Previous work shows that they process numbers as discrete tokens rather than quantities with inherent relationships,8,36 which may explain some of their limitations in numerical and decision-making contexts. Yet, our results suggest that LLMs can still form consistent relational boundaries in monetary settings, hinting at an implicit linguistic understanding of value. Second, we need to establish what behavior would be normatively appropriate, including how AI assistants should value comfort, fairness, privacy, and personalization in ways that align with user expectations and everyday decision making. Our findings show that variations in user background descriptions can lead to major outcome differences, mirroring previous work demonstrating sensitivity to user attributes such as gender, race, or dialect,19,24,25 and closer alignment with Western cultural norms.5,33 Third, we must determine how such alignment can be achieved through model design and communication. This includes studying how different inputs and personal data influence outcomes, and how interaction strategies such as presenting alternative options or brief clarifying questions can foster trust and better decisions. As LLMs become more involved in everyday choices, addressing these questions is essential to ensure their behavior aligns with human values and expectations.

Acknowledgments

We acknowledge the support of the “Onderzoeksprogramma Artificiële Intelligentie (AI) Vlaanderen” (FAIR), and the Research Foundation Flanders (FWO, grants G0G2721N and 1247125N).

    • 1. Acharya, D.B., Kuppan, K., and Divya, B. Agentic AI: Autonomous intelligence for complex goals - A comprehensive survey. IEEE Access (2025).

    • 4. Bhattacharya, R. and Aoun, M.A. Using generative AI in finance, and the lack of emergent behavior in LLMs. Commun. ACM 67, 8 (Aug. 2024), 67.

    • 5. Cao, Y., Zhou, L., and Lee, S. 
      Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study. In Proceedings of the FirstWorkshop on Cross-Cultural Considerations in NLP (C3NLP). Association for Computational Linguistics (2023), 5367.

    • 6. Chen, M., Tworek, J., and Jun, H. Evaluating large language models trained on code. arXiv preprint (2021); https://arxiv.org/pdf/2107.03374

    • 7. Chen, Y., Liu, T.X., and Shan, Y. The emergence of economic rationality of GPT. Proceedings of the National Academy of Sciences 120, 51 (2023).

    • 8. Davies, A.O., Nzoyem, R., and Ajmeri, N. Language models do not embed numbers continuously. arXiv preprint (2025), arXiv:2510.08009

    • 9. Deloitte. Autonomous generative AI agents: Under development. Deloitte Insights  (Jan. 2025).

    • 10. Denning, P.J. In large language models we trust?Commun. ACM 68, 6 (Jun. 2025), 2325.

    • 11. Dong, G., Wang, H., and Sun, J. Evaluating and mitigating linguistic discrimination in large language models. arXiv preprint (2024); https://arxiv.org/pdf/2404.18534

    • 12. Feng, G., Zhang, B., and Gu, Y. Towards revealing the mystery behind chain of thought: a theoretical perspective. Advances in Neural Information Processing Systems 36, (2023), 7075770798.

    • 14. Goethals, S. and Rhue, L. One world, one opinion? The superstar effect in LLM responses. In Proceedings of the 3rd Workshop on Cross-Cultural Considerations in NLP. Association for Computational Linguistics (2025), 89107.

    • 15. Goli, A. and Singh, A. Frontiers: Can large language models capture human preferences? Marketing Science 43, 4 (2024), 709722.

    • 16. Grattafiori, A., Dubey, A., and Jauhri, A. The Llama 3 herd of models. arXiv preprint (2024); https://arxiv.org/pdf/2407.21783

    • 17. Greengard, S. Is it possible to truly understand performance in LLMs?Commun. ACM 67, 12 (Nov. 2024), 1416.

    • 18. Hendrycks, D., Burns, C., and Basart, S. Measuring massive multitask language understanding. arXiv preprint (2020); https://arxiv.org/pdf/2009.03300

    • 19. Hofmann, V., Kalluri, P.R., and Jurafsky, D. AI generates covertly racist decisions about people based on their dialect. Nature 633, 8028 (2024), 147154.

    • 20. Jia, J.J., Yuan, Z., and Pan, J. Decision-making behavior evaluation framework for LLMs under uncertain context. Advances in Neural Information Processing Systems 37 (2024), 113360113382.

    • 22. Kamins, M.A., Folkes, V.S., and Fedorikhin, A. Promotional bundles and consumers' price judgments: When the best things in life are not free. J. of Consumer Research 36, 4 (Dec. 2009), 660670.

    • 23. Keeling, G., Street, W., and Stachaczyk, M. Can LLMs make trade-offs involving stipulated pain and pleasure states? arXiv preprint (2024); https://arxiv.org/pdf/2411.02432

    • 24. Kotek, H., Dockum, R., and Sun, D. Gender bias and stereotypes in large language models. In Proceedings of the ACM Collective Intelligence Conf. ACM (2023), 1224.

    • 25. Liang, P.P., Wu, C., and Morency, L.P. Towards understanding and mitigating social biases in language models. In Intern. Conf. on Machine Learning. PMLR (2021),65656576.

    • 27. Liu, R., Geng, J., and Peterson, J. Large language models assume people are more rational than we really are. In The Thirteenth Intern. Conf. on Learning Representations (2025).

    • 28. Mitchell, M., Attanasio, G., and Baldini, I. SHADES: Towards a multilingual assessment of stereotypes in large language models. In Proceedings of the 2025 Conf. of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1. Association for Computational Linguistics (2025), 1199512041.

    • 29. Ouyang, S., Zhang, J.M., and Harman, M. An empirical study of the non-determinism of ChatGPT in code generation. ACM Trans. on Software Engineering and Methodology 34, 2 (2025), 128.

    • 30. Purdy, M. What is agentic AI, and how will it change work? Harvard Business Rev (2024).

    • 31. Raman, N., Lundy, T., and Amouyal, S. STEER: Assessing the economic rationality of large language models. arXiv preprint (2024); https://arxiv.org/pdf/2402.09552

    • 32. Ross, J., Kim, Y., and Lo, A.W. LLM economicus? Mapping the behavioral biases of LLMs via utility theory. arXiv preprint (2024); https://arxiv.org/pdf/2408.02784

    • 33. Tao, Y., Viberg, O., and Baker, R.S. Cultural bias and cultural alignment of large language models. PNAS Nexus 3, 9 (2024), 346.

    • 34. Team, G.et al. Gemini: A family of highly capable multimodal models. arXiv preprint (2023), arXiv:2312.11805

    • 35. Vonasch, A.J., Mofradidoost, R., and Gray, K. People reject free money and cheap deals because they infer phantom costs. Personality and Social Psychology Bulletin (2024).

    • 36. Wallace, E., Wang, Y., and Li, S. Do NLP models know numbers? Probing numeracy in embeddings. arXiv preprint (2019), arXiv:1909.07940

    • 37. Wei, J., Wang, X., and Schuurmans, D. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 2482424837.

    • 38. Whiting, K. The rise of ‘AI agents’: What they are and how to manage the risks. World Economic Forum (Dec. 2024).

    • 39. Zao-Sanders, M. How people are really using gen AI in 2025. Harvard Business Rev. (Apr. 2025).

    • 40. Zhong, Q., Yun, Y., and Sun, A. Cultural value differences of LLMs: prompt, language, and model size. arXiv preprint (2024); https://arxiv.org/pdf/2407.16891

Mateusz Cedro (mateusz.cedro@uantwerpen.be) is a doctoral student at the University of Antwerp, Antwerp, Flanders, Belgium.

Timour Ichmoukhamedov is a postdoctoral researcher at the University of Antwerp, Antwerp, Flanders, Belgium.

Sofie Goethals is an assistant professor at the University of Antwerp, Antwerp, Flanders, Belgium.

Yifan He is a doctoral student at the University of Antwerp, Antwerp, Flanders, Belgium.

James Hinns is a doctoral student at the University of Antwerp, Antwerp, Flanders, Belgium.

David Martens is a professor at the University of Antwerp, Antwerp, Flanders, Belgium.

Submit an Article to CACM

CACM welcomes unsolicited submissions on topics of relevance and value to the computing community.

You Just Read

Cash or Comfort? How LLMs Value Your Inconvenience

View in the ACM Digital Library
This work is licensed under a Creative Commons Attribution International 4.0 license.
© 2026 Copyright held by the owner/author(s).