Overview of output performance
| Variable | Variant | Correct outcome in (%) |
|---|---|---|
| GPT model | ChatGPT 3.5 | 26.3 |
| ChatGPT 4.0 | 46.3 | |
| ChatGPT 4o | 52.5 | |
| ChatGPT o1 preview | 65.0 | |
| ChatGPT 4o mini | 28.8 | |
| ChatGPT o1 mini | 25.0 | |
| Task category* | Task A: Requesting specific legal references | 48.8 |
| Task B: Identification of legal references in context | 47.5 | |
| Task C: Case-based calculation | 28.8 | |
| Task D: Case-based decision making | 65.0 | |
| Prompting techniqure* | Zero-shot prompt | 45.0 |
| Few-shot prompt | 50.0 | |
| Accounting standard* | Local GAAP (Austrian GAAP) | 34.4 |
| International financial reporting standards (IFRS) | 60.6 |
| Variable | Variant | Correct outcome in (%) |
|---|---|---|
| ChatGPT 3.5 | 26.3 | |
| ChatGPT 4.0 | 46.3 | |
| ChatGPT 4o | 52.5 | |
| ChatGPT o1 preview | 65.0 | |
| ChatGPT 4o mini | 28.8 | |
| ChatGPT o1 mini | 25.0 | |
| Task category* | Task A: Requesting specific legal references | 48.8 |
| Task B: Identification of legal references in context | 47.5 | |
| Task C: Case-based calculation | 28.8 | |
| Task D: Case-based decision making | 65.0 | |
| Prompting techniqure* | Zero-shot prompt | 45.0 |
| Few-shot prompt | 50.0 | |
| Accounting standard* | Local | 34.4 |
| International financial reporting standards ( | 60.6 |
*Results without “mini”-variants
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.