Benchmark-Driven Evaluation Frameworks for AI Agent Trustworthiness in Enterprise Analytics
Keywords:
AI agent trustworthiness; enterprise analytics; text-to-SQL evaluation; benchmark design; LLM-as-judge; responsible AI adoptionAbstract
Enterprise adoption of AI-powered analytics agents has outpaced the methods available to evaluate whether their outputs can be trusted. General-purpose benchmarks such as MMLU [1] and HumanEval [2] measure static knowledge recall and isolated code-generation ability, while text-to-SQL benchmarks such as Spider [3] and BIRD [4] test query generation against simplified schemas. None of these evaluate the full chain of reasoning an enterprise analytics agent must perform: schema interpretation, organization-specific metric computation, calibrated uncertainty, and application of the correct analytical methodology. On BIRD, even the strongest reported system reaches only 40.08% execution accuracy against realistic enterprise schemas, compared with 92.96% for human analysts [4], a gap that becomes consequential once an agent's output informs financial or regulatory decisions rather than exploratory analysis. This article proposes a nine-dimension trustworthiness evaluation framework purpose-built for enterprise analytics agents. Four dimensions are grounded in established trust taxonomies [6]-[8]; five are analytics-specific contributions not addressed by any existing framework: Schema Fidelity, Metric Definition Adherence, Failure Mode Transparency, Business Context Understanding, and Analytical Methodology Competence. The article also proposes a benchmark design methodology combining automated, LLM-as-judge, and domain-expert graders, and a five-phase Continuous Benchmark Evolution Lifecycle for keeping evaluation infrastructure current as schemas, metric definitions, and agent capabilities change.
Downloads
References
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, "Measuring Massive Multitask Language Understanding," in Proc. ICLR, 2021. arXiv:2009.03300.
M. Chen, J. Tworek, H. Jun, et al., "Evaluating Large Language Models Trained on Code," arXiv:2107.03374, 2021.
T. Yu, R. Zhang, K. Yang, et al., "Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task," in Proc. EMNLP, 2018. DOI: 10.18653/v1/D18-1425.
J. Li, B. Hui, G. Qu, et al., "Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQL Evaluation," in Proc. NeurIPS Datasets and Benchmarks Track, 2023. arXiv:2305.03111.
P. Liang, R. Bommasani, et al., "Holistic Evaluation of Language Models," arXiv:2211.09110, 2022.
National Institute of Standards and Technology, "AI Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, 2023. DOI: 10.6028/NIST.AI.100-1.
B. Wang, W. Chen, H. Pei, et al., "DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models," in Proc. NeurIPS, 2023. arXiv:2306.11698.
Y. Huang, L. Sun, H. Wang, et al., "TrustLLM: Trustworthiness in Large Language Models," arXiv:2401.05561, 2024.
Accenture, "From Compliance to Confidence: Embracing a New Mindset to Advance Responsible AI Maturity," Accenture Research, 2024.
Accenture, "Technology Vision 2025," Accenture Research, 2025.
L. Zheng, W. Chiang, Y. Sheng, et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," in Proc. NeurIPS Datasets and Benchmarks Track, 2023. arXiv:2306.05685.
X. Liu, H. Yu, H. Zhang, et al., "AgentBench: Evaluating LLMs as Agents," in Proc. ICLR, 2024. arXiv:2308.03688.
S. Chang, J. Wang, M. Dong, et al., "Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness," in Proc. ICLR, 2023. arXiv:2301.08881.
D. Gao, H. Wang, Y. Li, et al., "Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation," arXiv:2308.15363, 2023.
X. Wang, Z. Wang, J. Liu, et al., "MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback," in Proc. ICLR, 2024. arXiv:2309.10691.
M. Pourreza and D. Rafiei, "DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction," in Proc. NeurIPS, 2023. arXiv:2304.11015.
Y. Lai, C. Li, Y. Wang, et al., "DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation," in Proc. ICML, 2023. arXiv:2211.11501.
L. Huang, W. Yu, W. Ma, et al., "A Survey on Hallucination in Large Language Models," arXiv:2311.05232, 2023.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All papers should be submitted electronically. All submitted manuscripts must be original work that is not under submission at another journal or under consideration for publication in another form, such as a monograph or chapter of a book. Authors of submitted papers are obligated not to submit their paper for publication elsewhere until an editorial decision is rendered on their submission. Further, authors of accepted papers are prohibited from publishing the results in other publications that appear before the paper is published in the Journal unless they receive approval for doing so from the Editor-In-Chief.
IJISAE open access articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. This license lets the audience to give appropriate credit, provide a link to the license, and indicate if changes were made and if they remix, transform, or build upon the material, they must distribute contributions under the same license as the original.


