The study aims to analyze ChatGPT’s performance in financial accounting judgment tasks, focusing on variations by task category, model version, prompting method and accounting standards.
The authors conducted an experimental design using 480 prompts across four task categories, six GPT models, two prompting techniques and two accounting standards [International Financial Reporting Standards (IFRS) and Austrian generally accepted accounting principles (GAAP)]. Experts rated outputs, which were analyzed using logistic regression and predictive machine learning.
The results indicate that the GPT model has a strong influence on whether the response is correct, with newer models performing better. The task category is a major factor, with GPT models struggling especially with rule-based calculations that are straightforward for human professionals. Performance is significantly lower for Austrian GAAP compared to IFRS, likely due to limited training data. Prompting has a limited overall impact.
The findings reveal ChatGPT’s potential and limits in accounting judgment, implying research needs for task delegation and the integration of frameworks.
While ChatGPT demonstrates potential as a digital assistant for specific tasks (with an accuracy up to 77%), expert oversight remains critical to ensure output performance. The presence of false positives can create a misleading sense of security, underscoring the need for professional judgment in compliance-sensitive contexts.
This study introduces a novel task taxonomy based on the Dreyfus model to assess ChatGPT’s performance in accounting judgment. It uniquely combines task complexity, accounting standards and model architecture to identify key performance drivers and frames ChatGPT through the lens of functional agency, offering a new perspective on its professional applicability in financial accounting judgment.
1. Introduction
In today’s dynamic economic environment, and due to increasingly complex and digital business models (Rieg and Radeck, 2024), the discipline of accounting encounters far-reaching changes. These changes fester in the form of an increase in required internal and external data sources for reporting due to the rapidly evolving information demands of relevant stakeholders, including investors, banks and national and international standard-setting bodies. Moreover, standard setters are adapting existing accounting standards and reporting requirements to the dynamic environment [e.g. accounting for digital assets (Jackson and Luu, 2023; Leitner-Hanetseder and Lehner, 2022)]. Consequently, the number and scope of accounting standards that a company must comply with continue to rise constantly. Furthermore, the pace of updates is high, especially regarding the imposed regulations in the context of sustainability accounting (Baumüller and Grbenic, 2021). Ensuring compliance with evolving accounting standards has become an increasingly intricate task (Knudsen, 2020), particularly for small and medium-sized entities (Malo et al., 2024).
As a result, the field of accounting has shifted to a discipline of highly specialized experts, which shows considerable implications for the department’s required staff size and the qualifications necessary to work within the department (Pargmann et al., 2023). One proposed solution to mitigate the high need for specialization and personnel power is using automation technology such as RPA (robotic process automation) and AI (artificial intelligence), especially the two sub-areas ML (machine learning), as well as GenAI (generative artificial intelligence) (Leitner-Hanetseder et al., 2021; Perdana et al., 2023).
However, while research on RPA use in accounting is quite advanced [e.g. to automate repetitive, rule-based tasks such as data reconciliation, invoice processing, and transaction posting (Schulze and Nuhn, 2020; Perdana et al., 2023)], there seems to be a lack of research on the application of GenAI and the use of LLMs (Large Language Models), particularly in the context of evaluating the output performance of financial accounting judgment simulated by LLMs.
LLMs represent a significant advancement in artificial intelligence, designed to understand, generate, and process human language. Trained on vast amounts of textual data, these models employ deep learning techniques, specifically neural networks, to analyze and predict language patterns (Le Duong et al., 2024). Among these models, GPT (Generative Pre-trained Transformer) is promising for its broad application capabilities, including analyzing financial reports, generating and summarizing texts, responding to queries on accounting policies, or interpreting accounting regulations, and therefore solving complex problems across the field of accounting (Dong et al., 2024; Zhao and Wang, 2024). It is therefore of practical and academic interest to examine the extent to which different types of tasks in the field of accounting judgment can be answered by LLMs correctly to comply with accounting standards.
In this paper, we add to that discussion and focus on the question of the extent to which accountants might be supported by the chatbot “ChatGPT” and its GPT models. We particularly want to investigate the role that GPT models can play in assisting accountants with questions related to the application of accounting standards. We identify and test four different tasks relevant to accounting, namely (A) Requests for specific legal references; (B) Identification of legal references within context; (C) Case-based calculations; and (D) Case-based decision-making. These tasks represent different levels of complexity, reflecting the diverse analytical and decision-making demands within financial accounting practice.
Furthermore, we aim to evaluate the influence of prompting techniques, which involve crafting specific instructions or examples to guide AI responses, such as zero-shot (answering without prior examples) and few-shot prompting (providing examples to guide responses), on the correctness of results. In addition, we conduct a comparative analysis of the search results to evaluate whether the GPT models yield better results for issues relating to the Austrian Commercial Code [Austrian local generally accepted accounting principles (GAAP)] than to International Financial Reporting Standards (IFRS).
The rest of the paper is structured as follows: Section 2 provides background information on LLMs, explores their applications in financial accounting, identifies accounting tasks in the field of accounting judgment, and presents an agency framework as well as related hypotheses to assess the gathered empirical data. Section 3 describes the scope and the methodology of the empirical research, and is followed by section 4, presenting the results. Section 5 draws main conclusions and provides insights into limitations and suggestions for future research.
2. Background
2.1 Large language models and ChatGPT
LLMs are machine learning models that can understand and generate natural language (Dong et al., 2024) and are trained on huge amounts of text to understand the details, patterns, and grammar of human language (Brown, 2020). The most well-known example is GPT by OpenAI, which can answer questions and fulfill requests. LLMs, particularly models based on the Generative Pre-trained Transformer (GPT) approach, generate each next token in a text one by one, predicting the next part of the text as they go. This “generative approach” distinguishes them from the older models. The process starts with an input text (“prompt”), and the model generates an output by sequentially predicting each next token. The scalability of these models and the diversity of their training data enable advanced language-processing capabilities. Different GPT models such as GPT 4o, GPT 4, GPT 3.5, GPT o1 preview, Claude, or Code Llama, vary in parameter size, fine-tuning, and availability (de Kok, 2024).
The ChatGPT 3.5 series, released in November 2022, introduced essential improvements in coherence and context comprehension compared to GPT 3. Models in this series were optimized for handling longer, more complex texts and had a refined sense of tone and style (Brown, 2020). GPT 4, launched in March 2023, brought a much larger context window of 8,192 tokens and around 1.7 trillion parameters, enabling it to understand extended, context-rich texts. For the first time, GPT 4 could also process images and speech as input and access web content when given a URL, expanding the model’s abilities beyond purely text-based tasks (OpenAI, 2023). This multimodal capability opened new possibilities for the use across different input types. In May 2024, OpenAI introduced GPT 4o, with “o” standing for “omni,” which represents the next level of multimodal functionality. GPT 4o processes text, image, and audio seamlessly, allowing for natural interactions across various input types. With an even larger context window and a very high token count, GPT 4o can handle complex requests and intelligently integrate information from different data sources (Dong et al., 2024). From ChatGPT 3.5 to GPT 4o, the token count in training progressively increased, enabling deeper language processing and better alignment with specific queries.
In July 2024, GPT 4o mini was introduced, offering a more cost-effective solution for applications such as customer support chatbots, and real-time data processing where cost and speed are critical factors (OpenAI, 2024a). GPT 4o mini is increasingly flexible and comprehensive, suitable for diverse applications ranging from text-based requests to complex, multimodal tasks. Furthermore, the o1 preview model from OpenAI, which was introduced in September 2024, showcases advanced reasoning capabilities, making it highly effective for complex tasks in science, coding, and mathematics. Unlike previous models, o1 preview is trained to spend more time reasoning through problems, which allows it to tackle intricate tasks with greater precision. Early tests have shown impressive results, with o1 preview achieving 83% accuracy on qualifying exams for the “International Mathematics Olympiad” (IMO) and ranking in the 89th percentile in Codeforces coding competitions (OpenAI, 2024b). This model is particularly valuable for researchers and professionals needing advanced problem-solving abilities, such as generating complex formulas or building multi-step workflows (OpenAI, 2024b). While o1 preview demonstrates superior capabilities in deep reasoning and possesses a comprehensive knowledge base, o1 mini is optimized for speed and efficiency, particularly suited for tasks related to science, technology, engineering, and mathematics (OpenAI, 2024c).
Next to the evolution in model capabilities, there exists a fundamental distinction between GPT and ChatGPT: While GPT refers to the underlying model trained on extensive text data with general language processing capabilities, ChatGPT is a specialized application fine-tuned to function as an interactive digital assistant (Mpofu et al., 2024; Lin et al., 2023). ChatGPT is specifically trained to provide human-like responses, conduct conversations, and follow user-friendly instructions (OpenAI, 2024d; Mpofu et al., 2024). The high processing capabilities, the translation of complex and domain specific language into something understandable, and easy access makes ChatGPT worthwhile investigating its applicability in accounting.
2.2 Applications and research of large language models in accounting
As mentioned, ChatGPT is a digital assistant that is used for tasks such as text generation (Wu et al., 2023; Mulia et al., 2023), code creation (Buscemi, 2023; Liu et al., 2023), knowledge querying (Wood et al., 2023; Zhang, 2023), and analytical tasks (Kim et al., 2024). Dong et al. (2024) identified in particular three key areas within financial accounting that deal with the application of LLM such as ChatGPT:
practical applications of LLMs in accounting;
implications on the roles of professionals and organizations; and
educational purposes.
Applications used in financial accounting support data processing in the background, for example, by creating complex macros in Excel (Fidel, 2023), serve as text-to-code generators for SQL queries and Python code, thereby reducing dependence on the IT department but enhancing the need for data and technology literacy. Additional applications of LLMs in the field of financial accounting discussed in the literature include extracting novel information from structured and unstructured sources (Li et al., 2024a; Moreno and Caminero, 2023), generating financial disclosures (Föhr et al., 2023; de Villiers et al., 2024), and updating accounting measurements. Furthermore, LLMs are integrated into robotic process automation applications to streamline the preparation of financial statements (Beerbaum, 2023). In addition, LLMs offer opportunities to streamline customer interactions (Rane, 2023; Zhao and Wang, 2024) and to support analysis and the detection of patterns and anomalies in financial statements (Rane, 2023), helping stakeholders to make more informed decisions (Kim et al., 2024; Singh and Singh, 2023).
Being available 24/7, digital assistants such as ChatGPT provide efficient, personalized learning at low cost, potentially expanding access to education and training opportunities for a wider audience (Vasarhelyi et al., 2023). The accounting education literature primarily explores LLMs capacity to solve professional accounting exams (Abeysekera, 2024; Albuquerque and Gomes dos Santos, 2024a; Atanasovski et al., 2023) or financial business cases (Cheng et al., 2023) to support students in their learning process. Furthermore, it evaluates whether LLMs can outperform human candidates in (professional) accounting exams (Atanasovski et al., 2023; Wood et al., 2023). Research suggests that AI-powered language models, like those used as accounting tutors, could significantly improve the efficiency of (re-)training accountants (Johnson et al., 2009). These AI-accounting tutors offer personalized learning experiences, adapting to the individual’s needs and pace. This approach is particularly valuable in helping professionals stay up to date with constantly evolving regulatory standards. Following Bloom’s idea that one-on-one tutoring significantly boosts student performance (Bloom, 1984), AI-accounting tutors (Johnson et al., 2009), despite lacking physical presence, could provide similar benefits (Vasarhelyi et al., 2023).
While interest in LLM’s capabilities continues to grow, its effectiveness in tasks requiring nuanced professional expertise – such as financial accounting judgment – remains underexplored. Notably, a study by Albuquerque and Gomes dos Santos (2024b) provides rare insights by examining LLM’s accuracy on recognition criteria for provisions under IAS 37. Such findings highlight the critical need to scrutinize LLMs not just as linguistic tools, but as decision-support systems and digital assistants, whose functional reliability must be rigorously assessed in domain-specific applications like financial accounting judgment.
2.3 Theoretical framework and hypotheses development
This study examines ChatGPT’s functional performance as a digital assistant in the context of professional accounting judgment. It builds on the agency framework proposed by Abeysekera (2024), who conceptualized ChatGPT in academic environments through three interconnected forms of agency: social, moral, and human agency. In this model, social agency captures the tool’s role as an interactive partner in knowledge construction, governed by norms, institutional roles, and user expectations. Moral agency addresses the ethical tensions arising from ChatGPT’s utilitarian knowledge distribution logic versus the deontological responsibility frameworks upheld in educational systems. Human agency, in turn, reflects the deliberate actions of individuals or institutions in shaping the responsible use of AI through intentionality, foresight, self-regulation, and reflection.
Transferring this model to the domain of professional financial accounting judgment requires a shift in focus: here, the central concern is no longer the possibilities and usage in the context of education and training, but output performance measured by correctness. While social and moral agency still play a role in shaping professional interaction and accountability, the decisive construct shifts toward functional agency. For a digital assistant to be meaningfully integrated into accounting workflows, it must deliver correct, compliant (with various accounting standards), and context-sensitive responses that meet the thresholds of professional correctness and legal accountability. To further explore what determines functional agency, we draw on the Technology Acceptance Model (TAM) (Davis, 1989). Within this framework, high output performance enhances both perceived usefulness and perceived ease of use, which are the key drivers of user acceptance. In a professional setting, an LLM such as ChatGPT will only be accepted as a digital assistant if it can reliably assist in judgment tasks, reduce cognitive load, and integrate seamlessly into the professional reasoning process. Thus, functional agency becomes the foundational form of agency in practice, shaping whether social or moral integration of the tool is even viable. Only when ChatGPT performs with sufficient accuracy can professionals reasonably begin to consider its integration as a communicative support tool (social agency) or reflect on its ethical implications (moral agency). This link is visualized in Figure 1.
The diagram illustrates financial accounting judgement in practice connected to moral and social agents, which influence a central human agent. The human agent is also linked with a functional agent, which directly connects to Chat G P T. Chat G P T links back to moral and social agents. Bidirectional arrows indicate interaction, suggesting dynamic influence between accounting practice, human decision-making, and A I tools such as Chat G P T.Adapted Agency Framework
Source: Framework by Abeysekera (2024) for financial accounting judgment
The diagram illustrates financial accounting judgement in practice connected to moral and social agents, which influence a central human agent. The human agent is also linked with a functional agent, which directly connects to Chat G P T. Chat G P T links back to moral and social agents. Bidirectional arrows indicate interaction, suggesting dynamic influence between accounting practice, human decision-making, and A I tools such as Chat G P T.Adapted Agency Framework
Source: Framework by Abeysekera (2024) for financial accounting judgment
The expanding role of ChatGPT in accounting offers new possibilities, but questions remain regarding its ability to provide correct answers and/or interpretations that align with accounting standards (Li and Vasarhelyi, 2024). Recent research has started to examine ChatGPT’s accuracy in addressing legal questions in the field of tax accounting (Alarie et al., 2023; Nay et al., 2023; Zhang, 2023). Both Alarie et al. (2023) and Zhang (2023) advise caution in relying on the responses of these models in complex or specific cases. Atanasovski et al. (2023) reported that ChatGPT 3.5 achieved a 73% pass rate on tasks in accounting exams. However, this pass rate must be interpreted with caution, as passing scores under typical grading schemes may still permit 40–50% incorrect answers.
Nevertheless, what becomes obvious is that output performance differences become evident across differences in tasks. More specifically, differences are expected to be significant between variations in categories such as rule-based recall or judgment-based reasoning. The findings are consistent with previous research in information technology (Hendrycks et al., 2021; Zhang et al., 2023) and accounting (Albuquerque and Gomes dos Santos, 2024a; Abeysekera, 2024; Atanasovski et al., 2023), leading us to hypothesize that there are also measurable differences in performance output when an AI agent (ChatGPT) is confronted with different task categories. Focusing on functional agency, task category serves as an independent variable, as it is systematically varied to assess its impact, while the correctness of the response of ChatGPT (=output performance) represents the dependent variable. Accordingly, we formulate the following hypothesis:
The output performance of accounting judgment generated by ChatGPT differs across task categories, reflecting systematic performance variation.
As outlined, OpenAI has developed a series of LLMs based on its GPT architecture (Abeysekera, 2024; Dong et al., 2024). ChatGPT refers to OpenAI’s chatbot application built on the GPT architecture (OpenAI, 2024d), with series 3.5 launched as an open-access version in November 2022 (Beerbaum, 2023). The subsequent model GPT 4o to GPT o1 preview mark significant advancements in language comprehension, contextual processing and multimodal capabilities. The empirical evidence suggests that earlier models, such as GPT 3.5, often struggle with complex accounting tasks in standardized exams (Abeysekera, 2024). In contrast, especially GPT 4o, achieves markedly higher output performance and has demonstrated the ability to pass various accounting assessments (Eulerich et al., 2023; Abeysekera, 2024). Based on research from other domains (Latif et al., 2024; de Winter et al., 2024), we anticipate that the GPT o1 model series, as a newer iteration specifically designed to handle system 2-like reasoning, which is inspired by the concept of Kaneman (2011) deliberate, analytical, and slow thinking, will likely build on this trend by further enhancing contextual understanding and accuracy in accounting.
On the other hand, mini models, developed for efficiency and speed, may not follow up on this upward performance trend. These models typically trade off complex reasoning capabilities for computational efficiency. For example, in the domain of academic statistics, McGee and Sadler (2025) report that while GPT 4o mini outperforms older models like GPT 3.5, it does not surpass the full GPT 4 model in output performance. Similarly, Wang and Wang (2025) show that GPT 4o mini does not consistently outperform GPT 3.5 or GPT 4o across financial and accounting tasks such as classification, sentiment analysis, summarization, text generation and prediction.
Despite these advancements, it remains uncertain whether newer models translate architectural improvements into correct outcomes in accounting judgment tasks. In this context, functional agency is also influenced by the specific model employed. The independent variable is the version of the GPT model used (e.g. GPT 3.5, GPT 4o, GPT o1 preview, and mini models), while the dependent variable is the output performance in accounting judgment tasks. This gap motivates the formulation of our second hypothesis:
Newer GPT model versions (e.g. GPT 4o and GPT o1 preview) achieve significantly higher output performance in accounting judgment tasks compared to earlier models (e.g. GPT 3.5), except for mini models, which may not show improvements with respect to output performance.
Accountants interact with ChatGPT by formulating prompts (Abeysekera, 2024). Prompts are instructions or questions designed to elicit a specific response or action from ChatGPT. These prompts often include contextual information and precise wording to enable a targeted and relevant answer. The structure and quality of the prompt directly impact the output performance and relevance of the generated response. Clear and specific prompts can significantly enhance output performance, as they help the model to precisely understand the context and intent of the question (Krause, 2023).
Current research on prompting techniques in accounting primarily focuses on financial analysis (Krause, 2023; Xie, 2024; Fatouros et al., 2023) or cost accounting (Xie, 2024) rather than on accounting judgments per se. Therefore, we see a need to examine how users’ specific prompting techniques impact the accuracy rate when asking questions dealing with accounting judgment. Several prompting techniques are currently discussed in the literature, with zero-shot and few-shot prompting being the most relevant options. Nay et al. (2023) indicate that few-shot prompting enhances the output performance of the GPT 4 model. Using zero-shot prompting, the model answers a question without any prior examples provided by the user, relying on its preexisting training alone. In contrast, few-shot prompting includes a few examples within the prompt, which often improves model output performance by setting clearer expectations for response format and context. Studies, like those from Krause (2023) or Nay et al. (2023), suggest that few-shot prompting significantly enhances results in more advanced models, such as GPT 4.
However, research also indicates that the effectiveness of prompting techniques may vary depending on the model architecture. For instance, Liu et al. (2024) found that advanced prompting techniques such as chain of thought can reduce accuracy in the o1 preview model compared to zero-shot prompting. Similarly, de Winter et al. (2024) demonstrated that o1 preview achieves near-perfect results on math exams without needing examples, pointing to the strength of zero-shot prompting when the model is well-aligned with systematic reasoning tasks. These findings suggest that the relationship between prompting technique (independent variable) and output performance (dependent variable) may vary significantly depending on the specific model and domain, thereby influencing the extent to which functional agency can be realized in practice. Given the critical role of prompting in shaping ChatGPT’s responses, it is essential to examine whether the use of example-based prompting affects the accuracy of tasks in accounting judgment. This leads to the following hypothesis:
Example-based prompting (e.g. few-shot) improves output performance in accounting judgment tasks compared to zero-shot prompting.
Addressing highly specialized (financial accounting) areas requires specific domain expertise. For LLMs to generate reliable answers, a sufficient and mostly accurate training data set is necessary. We assume that for Austrian local GAAP, the training data needed might often be too small in size, or similar to standards in other legislations, which, according to literature, might confuse the algorithm (Zhang, 2023; Abeysekera, 2024).
ChatGPT was trained with a wide range of general data that is not specifically tailored to individual accounting standards such as IFRS (International Financial Reporting Standards) or the Austrian Commercial Code. As a result, the quality of the output may differ depending on the accounting standard in question. To date, there has been no systematic analysis regarding whether and to what extent the quality of results varies for different accounting standards, such as IFRS and Austrian local GAAP. In this context, the independent variable is the type of accounting standard applied (IFRS vs Austrian local GAAP), while the dependent variable is the output performance of ChatGPT’s responses in accounting judgment tasks. These considerations lead to our fourth hypothesis:
Output performance in accounting judgment tasks is higher under internationally standardized frameworks like IFRS compared to local standards like Austrian local GAAP.
Figure 2 shows how the presented hypotheses are integrated into a conceptual framework, which is the basis for our quantitative assessment and further analysis.
The diagram shows four predictors: G P T model with four nominal levels, task category with four nominal levels, prompting with two nominal levels, and accounting standard with two nominal levels. Each connects to output performance, a binary outcome, through hypotheses H1, H2, H3, and H4. The model highlights structured testing of how these inputs affect binary performance results.Research model and null hypotheses
Source: Authors’ own work
The diagram shows four predictors: G P T model with four nominal levels, task category with four nominal levels, prompting with two nominal levels, and accounting standard with two nominal levels. Each connects to output performance, a binary outcome, through hypotheses H1, H2, H3, and H4. The model highlights structured testing of how these inputs affect binary performance results.Research model and null hypotheses
Source: Authors’ own work
3. Methodology
3.1 Research design
The study uses exploratory research, which is based on the output performance of interactions with six GPT models. In a situation similar to a laboratory experiment, we obtained data using multiple GPT models (4 original + 2 mini models), different task categories (4), different prompting techniques (2), and a different set of accounting standards (2) applied. This approach is ideal for isolating variables that are intertwined in natural settings, as researchers can control conditions, manipulate variables, and measure influencing factors, and enables to identify causal relationships and test conditions that are rare or nonexistent in natural environments (Kachelmeier and King, 2002; Libby et al., 2002) and enables a precise analysis of the impact of individual factors on output performance in the specific field of financial accounting judgment (Kachelmeier and King, 2002; Libby et al., 2002; Pankoff and Virgil, 1970; Reimsbach and de Meyst, 2021).
To generate data, we have defined a total of five distinct topic areas that are relevant under both IFRS and Austrian GAAP. These five topics include:
recognition and measurement of tangible assets, specifically land and buildings;
inventories;
trade receivables;
provisions; and
selected disclosure requirements related to the preparation of financial statement notes.
In total, we defined two iterations per topic and associated the task category, which will be presented in section 3.2.1. The five topics were selected based on their central relevance to financial accounting and their recurring presence in accounting textbooks (Doupnik and Perera, 2007; Melville, 2008; Alexander et al., 2020) as well as practical reporting obligations. This selection aligns with approaches in accounting education literature (Abeysekera, 2024). The aim was to ensure topic diversity in terms of content structure, required reasoning, and frequency of exposure in educational and regulatory materials.
None of the prompts used in the experiment were tested during the planning phase to maintain objectivity in creating these tasks and to test performance output using ChatGPT. This precaution ensured that the selection of prompts remained independent of the capabilities of the GPT models, thereby avoiding potential biases in the experimental results. The implementation phase was carried out in April and October 2024. To mitigate learning effects, each prompt was entered into a new chat session.
Given that Austrian GAAP is predominantly utilized by accountants in Austria who use the German language, all prompts in this study were composed in German. Consequently, the IFRS prompts were also formulated in German to ensure comparability and to eliminate language as a confounding variable. Following this design, we hope to ensure that any observed differences in performance output were attributable to the standards rather than linguistic variations in the prompts. A complete documentation of the used prompts and the corresponding output is available upon request from the authors.
3.2 Independent variables
3.2.1 Task category.
In this paper, we investigate how GPT models could support accountants in applying accounting standards and making financial accounting judgments. While prior research, such as Atanasovski et al. (2023) and Abeysekera (2024), confirms that the output performance of language models tends to decline with increasing task complexity, the existing literature lacks a systematic categorization of accounting judgment tasks. Specifically, no study to date has aligned such task categories with expert-defined complexity benchmarks. This gap underscores the need for a structured taxonomy that distinguishes between tasks based on their inherent complexity and relevance to professional accounting judgment.
Therefore, we adopted an expert-driven approach for our task category classification. First, we conducted a focus group with five accounting experts – three academics and two practitioners – to identify tasks in accounting judgment. This focus group discussion led to the identification of four distinct categories of tasks that accountants may encounter in accounting judgment:
Category A: Requesting specific legal references;
Category B: Identification of legal references in context;
Category C: Case-based calculation; and
Category D: Case-based decision-making.
As such, an expert-driven approach may include subjective bias; we triangulated the findings for the tasks required in accounting textbooks to mitigate this concern. In conclusion, three of the four categories (category B, C and D) can also be derived explicitly from the requirements outlined in diverse financial accounting textbooks (Alexander et al., 2020; Doupnik and Perera, 2007; Melville, 2008).
Second, we aimed to define the complexity of a task from a human-centric perspective. To ensure a clear and consistent evaluation, we adopted a binary classification approach. Given that task complexity is a function of the accounting judgment itself (Bonner, 1994), certain tasks exhibit lower levels of complexity than others. However, some researchers posit that task complexity is relative to both the nature of the task and the individual performing it (Wood, 1986; Byström, 2002; Maynard and Hakel, 1997). Therefore, we examine task complexity from the perspective of a competent professional, as defined by the third level of the five-level skill acquisition model of Dreyfus’ (1980), which outlines five proficiency levels: novice, advanced beginner, competent, proficient and expert (Dreyfus, 1980). At the competent practitioner level in accounting, individuals progress from rule-based actions to goal-oriented problem-solving. Competent practitioners systematically analyze data, focus on critical elements to achieve specific outcomes, and integrate contextual knowledge to prioritize relevant factors. However, their approach remains deliberate rather than intuitive, relying on structured procedures rather than experience-driven judgment or holistic understanding (Dreyfus, 1980; Murphy and Hassall, 2020; Wood et al., 2023). To operationalize this classification, we applied the lens of a competent practitioner to evaluate whether each task category should be considered complex or noncomplex:
Noncomplex tasks are defined by clear procedural rules, low context dependency, minimal uncertainty and limited associated risk (McAulay et al., 1998; Murphy and Hassall, 2020). They typically involve routine applications of foundational knowledge, such as aligning accounting standards with well-defined transactions or performing calculations based on explicit data inputs. These tasks are structured, predictable, and do not require extensive professional judgment.
Complex tasks, by contrast, are characterized by high context sensitivity, ambiguity and the necessity to integrate multiple interdependent factors. They entail substantial uncertainty and risk, requiring advanced expertise, accumulated experience and the ability to evaluate nuanced and atypical scenarios such as the interpretation of extraordinary transactions or the resolution of conflicting legal and financial considerations. Such tasks align with the upper stages of Dreyfus’ (1980) model of skill acquisition (levels 4 and 5: proficient and expert) and demand deep professional judgment and strategic reasoning (McAulay et al., 1998; Murphy and Hassall, 2020).
Third, to gain a deeper understanding of the complexity of these categories, it is essential to evaluate whether the categories A-D are complex or noncomplex. Building on the definition of complex and noncomplex tasks from the point of view of a competent professional and in line with the study by Blayney et al. (2016), we asked the accounting experts from the focus group in a follow-up meeting to evaluate the complexity of each task, based on the third skill level “competent practitioners” of Dreyfus’ (1980) model of skill acquisition. Table 1 presents the assessment results, showing that only one of the four categories – Category D: Case-based decision making – was identified as complex from the perspective of a competent professional. In addition, the table provides detailed reasoning for the classification.
Task categories and complexity in financial accounting judgment*
| Task | Example | Level | Reasons |
|---|---|---|---|
| A: Requesting specific legal references | Which paragraph is relevant for the assessment of restructuring provisions? | Noncomplex |
|
| |||
| |||
| B: Legal reference identification in context | I am an accountant in an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15 in the following year, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. In the event of a cancelation a penalty of € 100,000 occurs. Which legal reference might be relevant for solving this matter? | Noncomplex |
|
| |||
| |||
| C: Case-based calculations | I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31 / 12)? Provide the calculation path. (Answer briefly) | Noncomplex |
|
| |||
| |||
| D: Case-based decision making | I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. In the event of a cancellation, a penalty of € 100,000 is incurred. Is the recognition of a provision required, and argue why? (Answer briefly) | Complex |
|
| Task | Example | Level | Reasons |
|---|---|---|---|
| A: Requesting specific legal references | Which paragraph is relevant for the assessment of restructuring provisions? | Noncomplex | For a competent professional, the task is straightforward, as it primarily involves identifying information from known sources |
The task is rule-based, with clear procedures, minimal judgment and no requirement for integrating multiple factors | |||
Requires familiarity with the legal standards, which is part of the core skill set of competent practitioners | |||
| B: Legal reference identification in context | I am an accountant in an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15 in the following year, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. In the event of a cancelation a penalty of € 100,000 occurs. Which legal reference might be relevant for solving this matter? | Noncomplex | For a competent professional, this task involves linking a situation to the correct legal reference, a process that is structured and within their domain expertise |
Context is provided clearly in the prompt, reducing ambiguity and enabling systematic application of knowledge judgment | |||
Does not require deep strategic judgment or advanced experience; it relies on structured problem-solving aligned with their skill set | |||
| C: Case-based calculations | I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31 / 12)? Provide the calculation path. (Answer briefly) | Noncomplex | For a competent professional, this task category involves a straightforward application of accounting standards combined with basic arithmetic calculations |
The nature of the task is structured, and it is clear what the input data is and what are expectations, making it manageable without requiring judgment beyond basic application of rules | |||
Competent practitioners are trained in combining conceptual understanding with numerical accuracy, and the task does not require higher-order problem-solving | |||
| D: Case-based decision making | I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. In the event of a cancellation, a penalty of € 100,000 is incurred. Is the recognition of a provision required, and argue why? (Answer briefly) | Complex | Involves applying accounting knowledge in a goal-oriented but context-dependent manner, pushing beyond structured rules |
*Original Prompts were written in German. The examples given were translated
Fourth, the four task categories (A-D) were classified as nominal variables, representing a nonhierarchical manipulation of task type. Each category encapsulates a distinct form of accounting judgment, systematically varied to observe differences in output performance. While these categories do not inherently encode complexity, we analyzed model output performance across them, to explore whether certain task categories are more challenging for GPT models. For instance, if accuracy is consistently higher in categories A-C but lower in category D, this pattern may suggest an implicit complexity gradient. To further investigate this relationship, we compared output performance by category with expert-assessed complexity ratings (complex vs noncomplex) derived using the Dreyfus model of skill acquisition. This allowed us to assess whether model performance trends align with professional judgments about task difficulty, even though task complexity was not part of the manipulated design.
3.2.2 GPT models.
The independent variable in this study is the version of the GPT model applied to accounting judgment tasks. Conceptually, it reflects the generational and architectural progression of OpenAI’s large language models. Operationally, we tested six distinct versions under identical conditions: GPT 3.5, GPT 4.0, GPT 4o mini, GPT o1 mini, GPT 4o, and GPT o1 preview. The models differ in architecture, parameter size, and optimization for reasoning tasks. Mini models, designed for speed and efficiency, generally offer reduced performance in tasks requiring deeper contextual analysis. For the purpose of analysis, the GPT models were treated as a nominal variable. Although the versions reflect a chronological sequence, the performance differences between them are not guaranteed to be linear or equidistant. Treating them nominally allows for an unbiased estimation of model-specific effects without imposing assumptions about interval-scale relationships between model generations.
3.2.3 Prompting technique.
Prompting style was varied using two common techniques: zero-shot (no examples provided) and few-shot (one illustrative example provided). Each prompt-task combination was administered under both techniques to test the effect of input conditioning on model output. This differentiation can be illustrated with the following example in Table 2.
Prompting examples*
| Zero-shot | Few-shot |
|---|---|
| Question: I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31 / 12)? Provide the calculation path. (Answer briefly) | How can the probability of a warranty claim outflow be calculated to recognize and assess a provision under IAS 37? |
| An entity sells goods with a warranty under which customers are covered for the cost of repairs of any manufacturing defects that become apparent within the first six months after purchase. If minor defects were detected in all products sold, repair costs of 1 million would result. If major defects were detected in all products sold, repair costs of 4 million would result. The entity’s past experience and future expectations indicate that, for the coming year, 75% of the goods sold will have no defects, 20% of the goods sold will have minor defects and 5% of the goods sold will have major defects. In accordance with IAS 37 paragraph 24, an entity assesses the probability of an outflow for the warranty obligations as a whole. The expected value of the cost of repairs is: (75% of nil) + (20% of 1 m) + (5% of 4 m) = 400,000 (IAS 37.24) | |
| Question: I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31/12)? Provide the calculation path. (Answer briefly) | |
| Question: In which sections of the Austrian Commercial Code is the valuation of inventories defined? (Please answer briefly in max. 2–3 lines) | Example: Which section of the Austrian Commercial Code defines the measurement of property, plant and equipment assets (e.g. production machinery)? – Property, plant and equipment are fixed assets (§ 198 Abs. 2). The measurement of fixed assets is defined in § 203. Depreciation of fixed assets is regulated in § 204. |
| Question: In which sections of the Austrian Commercial Code is the valuation of inventories defined? (Please answer briefly in max. 2–3 lines) |
| Zero-shot | Few-shot |
|---|---|
| Question: I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31 / 12)? Provide the calculation path. (Answer briefly) | How can the probability of a warranty claim outflow be calculated to recognize and assess a provision under |
| An entity sells goods with a warranty under which customers are covered for the cost of repairs of any manufacturing defects that become apparent within the first six months after purchase. If minor defects were detected in all products sold, repair costs of 1 million would result. If major defects were detected in all products sold, repair costs of 4 million would result. The entity’s past experience and future expectations indicate that, for the coming year, 75% of the goods sold will have no defects, 20% of the goods sold will have minor defects and 5% of the goods sold will have major defects. In accordance with | |
| Question: I am an accountant in the accounting department of an Austrian bicycle wholesaler. The balance sheet date is December 31. On November 15, we signed a sales contract with our customer Alpha. It states that on March 15, 2000 bicycles are to be delivered by us at €150 per unit. On December 20, we plan to order the relevant bicycles from our manufacturer and learn that, due to a price increase, they now cost €170 per unit. Sales costs amount to €5 per unit. A contract withdrawal is possible at €50 per bicycle and shall be executed if it helps to minimize the damage. A provision is required, what′s the amount recognized at year end (31/12)? Provide the calculation path. (Answer briefly) | |
| Question: In which sections of the Austrian Commercial Code is the valuation of inventories defined? (Please answer briefly in max. 2–3 lines) | Example: Which section of the Austrian Commercial Code defines the measurement of property, plant and equipment assets (e.g. production machinery)? – Property, plant and equipment are fixed assets (§ 198 Abs. 2). The measurement of fixed assets is defined in § 203. Depreciation of fixed assets is regulated in § 204. |
| Question: In which sections of the Austrian Commercial Code is the valuation of inventories defined? (Please answer briefly in max. 2–3 lines) |
*Original Prompts were written in German. The examples given were translated
3.2.4 Accounting standards: international financial reporting standards versus Austrian GAAP.
In the study, the set of accounting standards used serves as an independent variable with nominal characteristics. The independent variable distinguishes between prompts referencing either the International Financial Reporting Standards (IFRS) or the Austrian Commercial Code (Austrian local GAAP).
3.3 Dependent variable: output performance
The dependent variable in this study is output performance, defined as the binary classification of each ChatGPT response as either fully correct or incorrect. A response was considered correct only if it was free from any false, misleading, or incomplete statements. This strict definition reflects the professional requirements of financial accounting, where precision is critical and even minor errors can have significant regulatory or financial consequences. Partially correct answers were therefore systematically classified as incorrect to ensure that the evaluation aligned with real-world standards of professional judgment.
The task categories included in the study permitted objective and unambiguous assessments. Therefore, none of the tasks involved interpretative cases that remain unresolved in literature or in professional accounting practice. All tasks were designed with clearly defined correct solutions grounded in applicable accounting standards and accounting literature. Each of the 480 outputs was independently assessed by two financial accounting professionals, each with more than 20 years of experience. All responses were benchmarked against predefined reference solutions. Ratings were subsequently compared and reconciled to resolve any discrepancies through consensus.
Inter-rater agreement was exceptionally high, as reflected in Cohen’s Kappa of 0.92, indicating near-perfect reliability. Full agreement was reached across all tasks in categories A and B, which required the identification of a specific legal reference. These cases were unambiguous and allowed for clear-cut professional judgments. In category C, a single disagreement arose concerning whether a numerically correct answer lacking an explanation of the calculation should be considered valid; after review, it was accepted as correct. In Category D, one response was initially rated as adequate by one evaluator but was ultimately judged as too vague and general to be practically useful. Although factually not incorrect, it was classified as incorrect due to its lack of professional relevance and actionable content.
3.4 Statistical modeling approach
To analyze the factors influencing output performance of ChatGPT-generated answers in financial accounting judgment, we applied a logistic regression model. This approach is appropriate given the binary nature of the dependent variable, which classified responses as either correct or incorrect based on predefined criteria.
The model included the following independent variables: task category, GPT model version, accounting standard, and prompting technique. For analysis, we used various Python libraries (pandas, NumPy, matplotlib, scikit-learn, statsmodels) and applied them using JupyterLab. The code for logistic regression (for explanation, see section 4.2) and machine learning (for prediction, see section 4.3) is included in the Supplementary Material.
The logistic regression model provided a structured basis for evaluating which task category, topic, GPT model, prompting technique, and accounting standard influence performance output the most. All nominal predictors were dummy coded with appropriate reference categories. Diagnostic procedures included checks for multicollinearity via variance inflation factors (VIF), as well as model performance assessments using pseudo-R-squared and Model χ2. On this basis, hypotheses assessment was performed.
To further strengthen the applicability of our research, we used the data to also train a predictive model using machine learning. Results obtained in the statistical analysis were considered and used in feature selection. To ensure the robustness and predictive validity of the ML algorithm, the data set was split into training and test sets using an 80:20 ratio. The model was estimated using maximum likelihood estimation based on the training data and validated using the holdout test set. Model performance is measured and documented using common metrics including accuracy, precision, recall and F1-score.
4. Results
The detailed analysis reveals a multitude of inaccuracies in ChatGPT’s answers. This section will analyze where they originate in detail. We also refer to Appendix for a detailed account of the results. The most obvious influence on the decline of output performance is observed by the use of local accounting standards in combination with specific task categories.
4.1 Descriptive statistics
In total, we analyzed 480 observations (prompts and output recorded). Results on output performance are presented in Table 3. The descriptive data shows a clear indication of superiority of newer models over the previous ones. In particular, the GPT o1 preview model demonstrated the highest overall correctness (65.0%), clearly outperforming earlier versions such as GPT 3.5 (26.3%) and GPT 4 (46.3%). This suggests a substantial improvement in legal and accounting task performance, likely due to architecture optimizations and updated training data. Interestingly, the GPT mini models exhibit a clear weakness in handling tasks within the domain of financial accounting, with their overall correctness across all questions remaining roughly on par with the older GPT-3.5 model. In sum, model performance correlates strongly with size and version, confirming that newer and larger models consistently deliver more accurate outputs in financial accounting judgment. Thus, in further statistical analysis only the four “full” variants of GPT models are investigated in detail while mini models are discarded.
Overview of output performance
| Variable | Variant | Correct outcome in (%) |
|---|---|---|
| GPT model | ChatGPT 3.5 | 26.3 |
| ChatGPT 4.0 | 46.3 | |
| ChatGPT 4o | 52.5 | |
| ChatGPT o1 preview | 65.0 | |
| ChatGPT 4o mini | 28.8 | |
| ChatGPT o1 mini | 25.0 | |
| Task category* | Task A: Requesting specific legal references | 48.8 |
| Task B: Identification of legal references in context | 47.5 | |
| Task C: Case-based calculation | 28.8 | |
| Task D: Case-based decision making | 65.0 | |
| Prompting techniqure* | Zero-shot prompt | 45.0 |
| Few-shot prompt | 50.0 | |
| Accounting standard* | Local GAAP (Austrian GAAP) | 34.4 |
| International financial reporting standards (IFRS) | 60.6 |
| Variable | Variant | Correct outcome in (%) |
|---|---|---|
| ChatGPT 3.5 | 26.3 | |
| ChatGPT 4.0 | 46.3 | |
| ChatGPT 4o | 52.5 | |
| ChatGPT o1 preview | 65.0 | |
| ChatGPT 4o mini | 28.8 | |
| ChatGPT o1 mini | 25.0 | |
| Task category* | Task A: Requesting specific legal references | 48.8 |
| Task B: Identification of legal references in context | 47.5 | |
| Task C: Case-based calculation | 28.8 | |
| Task D: Case-based decision making | 65.0 | |
| Prompting techniqure* | Zero-shot prompt | 45.0 |
| Few-shot prompt | 50.0 | |
| Accounting standard* | Local | 34.4 |
| International financial reporting standards ( | 60.6 |
*Results without “mini”-variants
ChatGPT exhibits notably lower performance levels in certain areas of professional accounting, particularly where domain-specific rules must be applied accurately. While large language models demonstrate general reasoning capabilities, their handling of specialized tasks – especially those requiring precise regulatory logic – remains inconsistent. This becomes evident when analyzing the output performance across the defined task categories.
The use of the task categories A and B appears to present a similar level of challenge for ChatGPT. Task D, while frequently considered cognitively demanding in accounting practice, showed the highest model success rate. This is reflected in the output performance: tasks in category A and B were answered correctly in 48.8% and 47.5% of cases, respectively, while Task D reached 65.0%. Task C, involving case-based numerical calculations, had the lowest correct answer rate at just 28.8%.
Tasks like those in category C, although typically regarded as structured and straightforward by competent accounting professionals, were particularly problematic for GPT models. The issues observed were not primarily due to computational errors, but rather the result of incorrect application or omission of accounting principles – for example, failing to apply monthly straight-line depreciation correctly. This highlights a key limitation: the models struggle not with arithmetic per se, but with correctly interpreting and applying accounting principles in a consistent and context-aware manner.
Interestingly, in some cases (especially with task D) responses comprising superficial statements were more likely to be rated as correct than those providing detailed and specific explanations. As the test design did not require follow-up prompts or clarifications, superficial answers often met the minimum correctness threshold. However, their added value for accounting professionals is limited. This illustrates a broader concern as superficial correctness can mask deeper gaps in professional adequacy and tasks like category D can lead to inflated assessments of model competence when the quality of explanations is not rigorously evaluated.
Furthermore, prompting appears to have no significant impact on output performance (zero-shot: 45.0%, few-shot: 50.0%). However, when taking a closer look, zero-shot prompting shows inferior results for models without reasoning (zero-shot: 36.7%; few-shot: 46.7%), while it seems to be superior for the model variant including reasoning (zero-shot: 70.0%; few-shot: 60.0%).
Also, the accounting standard seems to influence output performance. For IFRS, where more text and examples are available online (larger audience, larger community, […]), results are superior compared to the niche standard Austrian GAAP. We see a correct answer rate of 60.6% for IFRS-related tasks, while performance dropped to 34.4% for Austrian GAAP. An analysis of ChatGPT’s output shows that references to Austrian GAAP mostly fall between § 192 and § 257, which cover accounting standards. However, irrelevant sections, such as § 244 et seq. of the Austrian GAAP, dealing with topics in the field of group accounting, are often included. Further, in Austrian GAAP, it becomes clear that GPT models cannot distinguish between superseded and current regulations. For example, § 205 Austrian GAAP is mentioned multiple times, which was removed in 2015. With respect to the format of the references, the quality varied significantly. Some were very detailed, citing exact paragraphs and lines, while others only provided broad section ranges. These findings indicate a general lack of model-specific training or exposure to Austrian GAAP, which negatively affects output performance for this standard. In the case of IFRS, the references are generally more precise and consistent. However, GPT 4o cited the IFRS Sustainability Standards in response to a prompt about disclosures in the notes, suggesting some random or inconsistent selections rather than a clear understanding of the specific accounting standard. Also, IFRS-related references were more consistent. However, in cases where multiple standards were cited, the range appeared somewhat arbitrary (e.g. “IAS 1.110–122”). The common abbreviation “et seq.” was not used by any model, but it could have significantly reduced the error rate.
4.2 Inferential statistics
To determine whether the findings from the previous section apply not only to our sample tasks but also to the broader task categories being tested, we performed statistical analysis. Given that the dependent variable is dichotomous (correct answer versus incorrect answer), a logistic regression is used. As discussed above, the independent variables task category, GPT model, prompting technique, and accounting standard are nominal variables. Additionally, we included the five applied areas (accounting topics) in which we operationalized our task categories. The Omnibus test of the logistic regression shows a significant increase (χ2 = 102.32; p < 0.01) when including the used predictors compared to its base model (LL-Null = –221.41, Log-Likelihood = –170.25). Results therefore indicate that the output performance can indeed be better predicted when the variables deemed relevant in this analysis are considered. With respect to the model’s power to explain variability in output accuracy, we report pseudo R2 of 0.2311. These values suggest a medium power; however, for assessing output performance, the values achieved by the presented model can be considered reasonably good in many practical applications, particularly for behavioral or social science data (Liu et al., 2023). Further details on coefficients, statistical significance, and odds ratios can be found in Table 4.
Logistic regression analysis
| Variable | Beta | Std. error | z ratio | p-value | Odds |
|---|---|---|---|---|---|
| Constant | −0.6886 | 0.4750 | −1.4500 | 0.1471 | 0.5023 |
| Task_B: Identification of legal references in context | −0.0661 | 0.3640 | −0.1820 | 0.8558 | 0.9360 |
| Task_C: Case-based calculation | −11.1680 | 0.3840 | −2.9100 | 0.0036 | 0.3273 |
| Task_D: Case-based decision-making | 0.8807 | 0.3740 | 2.3530 | 0.0186 | 2.4126 |
| Topic_Inventories | 0.3347 | 0.4100 | 0.8160 | 0.4143 | 1.3975 |
| Topic_Provisions | 13.6710 | 0.4280 | 3.1950 | 0.0014 | 3.9239 |
| Topic_Tangible assets | 0.4177 | 0.4100 | 1.0190 | 0.3084 | 1.5185 |
| Topic_Trade receivables | −0.7129 | 0.4270 | −1.6710 | 0.0946 | 0.4902 |
| GPT_Model_GPT 4.0 | 11.3300 | 0.3870 | 2.9260 | 0.0034 | 3.1051 |
| GPT_Model_GPT 4o | 14.5950 | 0.3900 | 3.7420 | 0.0002 | 4.3040 |
| GPT_Model_GPT o1 preview | 21.3250 | 0.4050 | 5.2640 | 0.0000 | 8.4360 |
| GAAP_Local GAAP | −14.1340 | 0.2750 | −5.1320 | 0.0000 | 0.2433 |
| Prompt_zero-shot | −0.2799 | 0.2650 | −1.0560 | 0.2912 | 0.7558 |
| LL-Null | −221,41 | ||||
| Log-Likelihood | −170,25 | ||||
| Model x² | 102,32 | ||||
| p | 0,000 | ||||
| Pseudo R² | 0,2311 | ||||
| n | 320 | ||||
| Variable | Beta | Std. error | z ratio | p-value | Odds |
|---|---|---|---|---|---|
| Constant | −0.6886 | 0.4750 | −1.4500 | 0.1471 | 0.5023 |
| Task_B: Identification of legal references in context | −0.0661 | 0.3640 | −0.1820 | 0.8558 | 0.9360 |
| Task_C: Case-based calculation | −11.1680 | 0.3840 | −2.9100 | 0.0036 | 0.3273 |
| Task_D: Case-based decision-making | 0.8807 | 0.3740 | 2.3530 | 0.0186 | 2.4126 |
| Topic_Inventories | 0.3347 | 0.4100 | 0.8160 | 0.4143 | 1.3975 |
| Topic_Provisions | 13.6710 | 0.4280 | 3.1950 | 0.0014 | 3.9239 |
| Topic_Tangible assets | 0.4177 | 0.4100 | 1.0190 | 0.3084 | 1.5185 |
| Topic_Trade receivables | −0.7129 | 0.4270 | −1.6710 | 0.0946 | 0.4902 |
| GPT_Model_GPT 4.0 | 11.3300 | 0.3870 | 2.9260 | 0.0034 | 3.1051 |
| GPT_Model_GPT 4o | 14.5950 | 0.3900 | 3.7420 | 0.0002 | 4.3040 |
| GPT_Model_GPT o1 preview | 21.3250 | 0.4050 | 5.2640 | 0.0000 | 8.4360 |
| GAAP_Local | −14.1340 | 0.2750 | −5.1320 | 0.0000 | 0.2433 |
| Prompt_zero-shot | −0.2799 | 0.2650 | −1.0560 | 0.2912 | 0.7558 |
| LL-Null | −221,41 | ||||
| Log-Likelihood | −170,25 | ||||
| Model x² | 102,32 | ||||
| p | 0,000 | ||||
| Pseudo R² | 0,2311 | ||||
| n | 320 | ||||
With respect to the main effects, GPT model (p < 0.001), task category (p < 0.001), and accounting standard (p < 0.001) all show significance. Only prompting shows no consistent effect on performance output. Therefore, our analysis leads to the acceptance of the formulated hypotheses:
H1, which hypothesized that the output performance of accounting judgment generated by ChatGPT differs across task categories, can be accepted.
H2, which hypothesized that newer GPT model versions (e.g. GPT 4o and GPT o1) achieve significantly higher output performance in accounting judgment tasks compared to earlier models (e.g. GPT 3.5), can be accepted.
H4, suggesting that output performance in accounting judgment tasks is higher under internationally standardized frameworks such as IFRS compared to local standards like Austrian GAAP, proves to be accurate and can be accepted.
While one hypothesis needs to be rejected:
H3, suggesting that example-based prompting (e.g. few-shot) improves output performance in accounting judgment tasks compared to zero-shot prompting, does not prove to be significant and needs to be rejected.
After creating separate dummies for each tested nominal topic, the effect sizes are also visible in Figure 3. The visualization clearly separates positive and negative influences and depicts confidence intervals at a 95% level for easy interpretation.
The graph plots logistic regression coefficients with 95 percent confidence intervals. Variables include G P T model versions, task categories such as case-based advice, calculation, and reference identification, as well as prompting type and accounting standard. Points left of zero indicate negative impact, while points to the right show positive influence. G P T Model G P T 01 and G P T 04 show strong positive coefficients near or above 2, while local G A A P shows a negative coefficient below negative 1.Overview results
Source: Authors’ own work
The graph plots logistic regression coefficients with 95 percent confidence intervals. Variables include G P T model versions, task categories such as case-based advice, calculation, and reference identification, as well as prompting type and accounting standard. Points left of zero indicate negative impact, while points to the right show positive influence. G P T Model G P T 01 and G P T 04 show strong positive coefficients near or above 2, while local G A A P shows a negative coefficient below negative 1.Overview results
Source: Authors’ own work
The analysis suggests that each increase in the variable GPT model increases the likelihood of the output being correct when confronted with accounting tasks. Regarding task category, a case-based decision-making question results in better output performance, while tasks requiring basic calculation drastically reduce output performance. With the prompting technique, output performance does neither increase nor decrease significantly. Finally, the availability of presumably more reference texts, as is the case for IFRS compared to Austrian local GAAP, increases output performance.
4.3 Predicting output performance using machine learning
Given our increased understanding of variables influencing output performance, we want to use the data also to predict future outcomes given practical tasks associated with the job of an accountant. We use all relevant variables from our statistical analysis, but render those irrelevant for feature engineering that did not prove significant (prompting, mini variants of the GPT models). For analysis, we have again 320 samples, which we separated into a data set for training (80%, n = 256) and testing (20%, n = 64). The logistic regression was trained using 1,000 iterations, and results are presented in the Table 5.
Logistic regression results (ML)
| Features | Model coefficients |
|---|---|
| Intercept | −0.0574 |
| Task_A: Legal reference | 0.1756 |
| Task_B: Identification of legal references in context | 0.0218 |
| Task_C: Case-based calculation | −0.9700 |
| Task_D: Case-based decision making | 0.7311 |
| Topic_Disclosure | −0.2664 |
| Topic_Inventories | −0.1370 |
| Topic_Provisions | 0.9577 |
| Topic_Tangible assets | 0.0336 |
| Topic_Trade receivables | −0.6294 |
| GPT_Model_GPT 3.5 | −1.0379 |
| GPT_Model_GPT 4.0 | −0.1271 |
| GPT_Model_GPT 4o | 0.2557 |
| GPT_Model_GPT o1 preview | 0.8678 |
| GAAP_IFRS | 0.5870 |
| GAAP_Local GAAP | −0.6285 |
| Model specifications | |
| Training N | 256 |
| Test N | 64 |
| f1-score | 0.77 |
| Precision | 0.77 |
| Recall | 0.77 |
| Features | Model coefficients |
|---|---|
| Intercept | −0.0574 |
| Task_A: Legal reference | 0.1756 |
| Task_B: Identification of legal references in context | 0.0218 |
| Task_C: Case-based calculation | −0.9700 |
| Task_D: Case-based decision making | 0.7311 |
| Topic_Disclosure | −0.2664 |
| Topic_Inventories | −0.1370 |
| Topic_Provisions | 0.9577 |
| Topic_Tangible assets | 0.0336 |
| Topic_Trade receivables | −0.6294 |
| GPT_Model_GPT 3.5 | −1.0379 |
| GPT_Model_GPT 4.0 | −0.1271 |
| GPT_Model_GPT 4o | 0.2557 |
| GPT_Model_GPT o1 preview | 0.8678 |
| GAAP_IFRS | 0.5870 |
| GAAP_Local | −0.6285 |
| Model specifications | |
| Training N | 256 |
| Test N | 64 |
| f1-score | 0.77 |
| Precision | 0.77 |
| Recall | 0.77 |
Model accuracy, precision, and recall are 77% which seems satisfying given the task of predicting output performance using a LLM. The confusion matrix, which is used to calculate recall and precision, shows seven false negatives (10.9%) and eight false positives (12.5%). The latter are the more concerning ones, as tasks that we might think we can solve using ChatGPT without problems show incorrect results. False positives may go unnoticed and lead to erroneous conclusions, especially in compliance-sensitive contexts such as financial accounting judgment.
5. Discussion
5.1 Conclusion and implications
The present study investigated the performance of ChatGPT across 480 accounting judgment tasks, covering four task categories, two accounting standards (IFRS and Austrian GAAP), two prompting techniques (few-shot and zero-shot prompting), and six GPT model versions. The results show clear performance differences depending on the task category, the accounting framework used and the GPT model employed:
The following findings emerged.
ChatGPT performed best in tasks requiring contextual interpretation and basic legal reference identification. In contrast, tasks involving case-based calculations, although noncomplex for competent professionals, frequently yielded incorrect answers due to misapplication of accounting logic.
Outputs related to IFRS were significantly more accurate than those under Austrian GAAP. This likely reflects the greater availability of IFRS training data and international relevance.
The most recent model (GPT o1 preview) showed the highest output performance across all categories. Mini models, in contrast, performed significantly worse, especially in complex decision-making or calculations.
Few-shot prompting improved output performance in earlier or less reasoning-optimized models. However, GPT o1 preview, which uses advanced reasoning, performed better with zero-shot prompting, suggesting that few-shot examples can reduce efficiency in models already optimized for reasoning.
Most importantly, the errors produced by ChatGPT are often subtle, not readily identifiable, and not confined to any one task category. This makes it particularly difficult to anticipate or detect when the model will fail. Even tasks that seem simple to humans can yield flawed results, often with high linguistic confidence. These inconsistencies undermine trust and present a serious obstacle to real-world deployment.
This insight is central to evaluating the potential of ChatGPT as a digital assistant in financial accounting judgment. While the idea of using LLMs to support or partially automate routine judgment tasks is appealing, current performance levels suggest that such tools are not yet ready for autonomous use. A digital assistant in accounting must be capable of applying regulatory knowledge accurately, interpreting context, and distinguishing between current and outdated standards and requirements. ChatGPT does not yet meet these requirements consistently.
At present, only the most advanced model variants begin to approach the level of functional agency required to support professional decision-making – and only under narrow conditions (e.g. IFRS-based, noncalculative tasks). Consequently, ChatGPT should be viewed not as a standalone assistant, but as a supplementary research tool. It may be useful in the early stages of standard identification or exploratory analysis, particularly when used under expert supervision.
Looking ahead, the steady improvements in GPT model performance suggest that the vision of a more capable digital assistant in accounting may become feasible. However, until the problem of confidently incorrect outputs is resolved, the responsibility for interpretation, compliance and final judgment must remain in human expert oversight (Leitner-Hanetseder et al., 2021).
Building on an adapted version of the agency framework proposed by Abeysekera (2024), we contend that without sufficient functional agency, the emergence of ChatGPT as a tool with social and/or moral agency in professional financial accounting remains premature. It cannot yet act as a reliable communicative partner (social agency) or take responsibility for the normative and professional reasoning required in accounting judgment (moral agency), because its performance outputs are too inconsistent and less robust. The path to trustworthy digital assistance in accounting is promising, but not yet complete.
5.2 Limitations
This study contributes valuable insights into the potential applications of ChatGPT and similar chatbots using LLMs in financial accounting, particularly in tasks involving accounting standards and financial judgment. However, several limitations must be acknowledged. These limitations are important and should be addressed in future research.
While the study investigated four representative task categories, these categories might not fully capture the entire spectrum of tasks accountants encounter in practice. However, the selected categories reflect key areas of financial accounting and offer a solid foundation for further exploration. Furthermore, the comparison between Austrian GAAP and IFRS represents only a subset of possible accounting standards. Results may vary significantly for other local GAAPs that were not included in this study.
In the field of industrial applications, an output accuracy level of 60% has been observed (Li et al., 2024b), while in the domain of accounting, to the best of our knowledge, no comparable benchmarks currently exist. However, the 60% threshold from industrial applications may not fully capture the requirements for critical accounting tasks, where even minor inaccuracies can lead to significant financial or legal consequences. Despite this, the threshold provides a practical starting point for assessing LLM performance in the absence of established accounting-specific benchmarks.
The study was conducted in a controlled environment, akin to a laboratory experiment. While this approach isolates influencing factors, it does not account for real-world complexities, such as integrating ChatGPT into existing accounting workflows or collaboration between human accountants and AI systems.
Responses were evaluated by human experts against predefined benchmarks, which, while ensuring professional oversight, may introduce subjective biases. Additionally, the binary assessment of correctness (correct/incorrect) does not account for partially correct answers that could be useful in practice. Nonetheless, this rigorous evaluation process mirrors the high accuracy standards required in accounting.
Given the rapid advancements in AI, the findings may become less relevant over time as newer, more advanced models are developed. This study focuses on models available by using ChatGPT as of October 2024, limiting its relevance as technology evolves.
Additionally, the study did not customize or optimize ChatGPT settings, such as temperature, a parameter which influences the randomness of the model’s responses. A lower temperature results in more deterministic and consistent answers, while a higher temperature introduces more creativity but may reduce accuracy (McTear and Ashurkina, 2024). Moreover, fine-tuning or other performance-enhancing techniques, such as providing domain-specific training data, were not applied. While these measures could potentially improve output performance, the study opted to evaluate the default configurations to maintain objectivity and ensure broad applicability of findings.
Another limitation is that the study did not examine whether repeated application of the same task would lead to different answers or improve output performance through potential learning effects or adaptations by the model. Such an investigation could provide valuable insights into ChatGPT’s ability to refine its outputs over time or its consistency when dealing with identical prompts.
Finally, the study did not explore ethical implications, such as the risk of overreliance on AI or the potential for errors in high-stakes situations. Similarly, regulatory concerns about the use of AI in financial accounting remain unaddressed. That said, the study underscores the importance of human expertise, mitigating ethical and compliance risks when AI is used as a supplementary tool.
5.3 Further research
Future research should explore the reasons underlying the rapid and high acceptance of ChatGPT among accountants (Ross and Zhang, 2024; Frenkenberger and Leitner-Hanetseder, 2024). An investigation of the factors contributing to this acceptance could yield valuable insights into the integration of the chatbot into accounting workflows. Investigating the factors influencing acceptance in ChatGPT’s capabilities in specific domains such as financial accounting judgment.
Additionally, examining methods to ensure quality assurance could significantly enhance the utility of ChatGPT as a digital assistant for tasks related to accounting standards and accounting judgment. This includes establishing clear frameworks for distinguishing between correct and incorrect responses and identifying which types of tasks can be reliably delegated to the assistant. A better understanding is required of what constitutes complexity and incorrect output from the model’s perspective and how this affects performance in domain-specific applications.
Furthermore, future research should explore methods to improve the traceability and transparency of AI-generated (Rane et al., 2023; Li et al., 2024a) such as ChatGPT’s responses, particularly by incorporating precise and authoritative source references. The ability for accountants to verify AI-generated outputs against external standards or regulations is a key factor in building trust and ensuring reliability. Therefore, the rapid adoption of AI tools like ChatGPT necessitates a reevaluation of the skills required by accountants (Boritz and Stratopoulos, 2023), ensuring they are equipped to effectively use AI tools while maintaining professional accounting judgment.
Also important is studying strategies for mitigating ethical and compliance risks associated with AI use in financial accounting (Khan and Umer, 2024; Sharma and Kavia, 2023). Transparency in decision-making processes, addressing biases embedded in training data (in accordance with the EU AI-act), and defining clear boundaries for AI applications are critical to maintaining professional accountability and adherence to regulatory standards. Establishing best practices for integrating AI such as LLMs into accounting workflows without compromising ethical responsibilities could enhance its practical value (Adeyeri, 2024).
Another promising area of research involves the technical integration of internal guidelines (Mathav et al., 2024). Internal guidelines, especially those that govern the application of accounting standards and options under regulatory frameworks, play a crucial role in accounting practice. Customizing LLMs to recognize and prioritize such internal guidelines would enable more tailored and context-sensitive applications. Therefore, research should address how LLMs can better reflect the hierarchical structure of legal and regulatory frameworks, ensuring that responses prioritize authoritative rules and properly navigate conflicting standards. This hierarchical awareness would be particularly useful in complex regulatory environments, where international standards, local laws, and organizational policies overlap.
By addressing these aspects, future research can contribute to a comprehensive framework for leveraging ChatGPT and similar tools in accounting. This framework would balance innovation with professional responsibility, ensuring that LLMs enhance efficiency and accuracy while maintaining compliance, ethical standards, and transparency.
Acknowledgements
This manuscript was reviewed and improved for language and grammar using the AI-based tool Paperpal.
References
Supplementary material
The supplementary material for this article can be found online
Appendix
Detailed results of output performance (Austrian local GAAP)
| Austrian GAAP | Category A | Category B | Category C | Category D |
|---|---|---|---|---|
| GPT 3.5 | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | False | False | False |
| Inventory | False | False | False | False |
| Receivables | False | False | False | False |
| Provisions | False | False | False | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | False | Correct |
| Provisions | False | False | False | Correct |
| Notes | False | False | False | Correct |
| GPT 4 | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | False | False | Correct | Correct |
| Provisions | False | False | Correct | Correct |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | Correct | False | False | Correct |
| Provisions | False | Correct | False | Correct |
| Notes | False | False | Correct | False |
| GPT 4o | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | False | False | False | Correct |
| Receivables | False | False | False | Correct |
| Provisions | Correct | Correct | Correct | False |
| Notes | False | False | Correct | False |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | False | Correct |
| Provisions | Correct | False | False | False |
| Notes | False | False | Correct | Correct |
| GPT 4o mini | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | Correct | Correct |
| Provisions | False | False | False | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | False | Correct | False | False |
| Provisions | False | False | False | False |
| Notes | False | False | False | Correct |
| GPT o1 preview | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | False | Correct | Correct |
| Receivables | Correct | False | Correct | Correct |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | False | False | Correct | Correct |
| Receivables | False | False | False | Correct |
| Provisions | Correct | False | Correct | False |
| Notes | Correct | False | False | False |
| GPT o1 mini | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | False | False |
| Provisions | False | False | Correct | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | Correct | Correct | False | False |
| Provisions | False | False | False | Correct |
| Notes | False | False | Correct | False |
| Austrian | Category A | Category B | Category C | Category D |
|---|---|---|---|---|
| Zero-Shot | ||||
| Fixed assets ( | False | False | False | False |
| Inventory | False | False | False | False |
| Receivables | False | False | False | False |
| Provisions | False | False | False | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | False | Correct |
| Provisions | False | False | False | Correct |
| Notes | False | False | False | Correct |
| Zero-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | False | False | Correct | Correct |
| Provisions | False | False | Correct | Correct |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | Correct | False | False | Correct |
| Provisions | False | Correct | False | Correct |
| Notes | False | False | Correct | False |
| Zero-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | False | False | False | Correct |
| Receivables | False | False | False | Correct |
| Provisions | Correct | Correct | Correct | False |
| Notes | False | False | Correct | False |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | False | Correct |
| Provisions | Correct | False | False | False |
| Notes | False | False | Correct | Correct |
| Zero-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | Correct | Correct |
| Provisions | False | False | False | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | False | Correct | False | False |
| Provisions | False | False | False | False |
| Notes | False | False | False | Correct |
| Zero-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | False | Correct | Correct |
| Receivables | Correct | False | Correct | Correct |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | False | False | Correct | Correct |
| Receivables | False | False | False | Correct |
| Provisions | Correct | False | Correct | False |
| Notes | Correct | False | False | False |
| Zero-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | False | False | Correct | False |
| Receivables | False | False | False | False |
| Provisions | False | False | Correct | False |
| Notes | False | False | False | False |
| Few-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | False | False | False | False |
| Receivables | Correct | Correct | False | False |
| Provisions | False | False | False | Correct |
| Notes | False | False | Correct | False |
Detailed results of output performance (IFRS)
| IFRS | Category A | Category B | Category C | Category D |
|---|---|---|---|---|
| GPT 3.5 | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | False | False |
| Provisions | False | Correct | False | Correct |
| Notes | Correct | False | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | False | False | False | False |
| Provisions | Correct | Correct | False | Correct |
| Notes | False | False | False | Correct |
| GPT 4 | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | Correct | Correct | False | False |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | False | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | False | Correct |
| Notes | Correct | False | False | False |
| GPT 4o | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | Correct | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | Correct | False | False | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | Correct | Correct |
| GPT 4o mini | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | False | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | Correct | False |
| Provisions | Correct | False | False | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | Correct | False | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | False | False |
| Provisions | Correct | False | False | Correct |
| Notes | Correct | False | False | Correct |
| GPT o1 preview | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | Correct | Correct |
| Receivables | False | Correct | False | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | Correct | Correct | False | Correct |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | False | Correct |
| GPT o1 mini | ||||
| Zero-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | False | False | False | Correct |
| Receivables | False | False | False | False |
| Provisions | False | False | False | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets (PPE) | False | Correct | False | Correct |
| Inventory | Correct | False | False | Correct |
| Receivables | False | False | False | False |
| Provisions | False | Correct | False | Correct |
| Notes | False | False | False | Correct |
| Category A | Category B | Category C | Category D | |
|---|---|---|---|---|
| Zero-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | False | False |
| Provisions | False | Correct | False | Correct |
| Notes | Correct | False | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | False | False | False | False |
| Provisions | Correct | Correct | False | Correct |
| Notes | False | False | False | Correct |
| Zero-Shot | ||||
| Fixed assets ( | Correct | Correct | False | False |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | False | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | False | Correct |
| Notes | Correct | False | False | False |
| Zero-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | Correct | False | False |
| Receivables | False | False | Correct | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | Correct | Correct |
| Few-Shot | ||||
| Fixed assets ( | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | Correct | False | False | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | Correct | Correct |
| Zero-Shot | ||||
| Fixed assets ( | False | False | False | Correct |
| Inventory | Correct | False | False | False |
| Receivables | False | False | Correct | False |
| Provisions | Correct | False | False | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | Correct | False | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | False | False | False | False |
| Provisions | Correct | False | False | Correct |
| Notes | Correct | False | False | Correct |
| Zero-Shot | ||||
| Fixed assets ( | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | Correct | Correct |
| Receivables | False | Correct | False | False |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | Correct | Correct | False | Correct |
| Inventory | Correct | Correct | False | Correct |
| Receivables | Correct | Correct | False | Correct |
| Provisions | Correct | Correct | Correct | Correct |
| Notes | Correct | Correct | False | Correct |
| Zero-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | False | False | False | Correct |
| Receivables | False | False | False | False |
| Provisions | False | False | False | Correct |
| Notes | False | False | False | Correct |
| Few-Shot | ||||
| Fixed assets ( | False | Correct | False | Correct |
| Inventory | Correct | False | False | Correct |
| Receivables | False | False | False | False |
| Provisions | False | Correct | False | Correct |
| Notes | False | False | False | Correct |

