FinTrust：金融领域可信度评估综合基准 (FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain)

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications. Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation. We assess eleven LLMs on FinTrust and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness. For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness. We believe that FinTrust can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.

翻译：近期，大型语言模型在解决金融相关问题方面展现出良好潜力。然而，由于金融领域的高风险与高利害属性，将大型语言模型应用于实际金融场景仍面临挑战。本文提出FinTrust——一个专门为评估大型语言模型在金融应用中的可信度而设计的综合基准。该基准基于实际应用场景，聚焦广泛的对齐问题，并为可信度评估的每个维度设计了细粒度任务。我们在FinTrust上评估了十一个大型语言模型，发现如o4-mini等专有模型在安全性等多数任务中表现更优，而如DeepSeek-V3等开源模型则在行业级公平性等特定领域具有优势。对于受托责任对齐与信息披露等挑战性任务，所有大型语言模型均表现不足，显示出其在法律意识方面存在显著差距。我们相信FinTrust能够成为金融领域大型语言模型可信度评估的重要基准。

相关内容

MoDELS

关注 44

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日