GPT作为知识工作者:对(AI)CPA能力进行零热评价 (GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities) - 专知论文

会员服务 ·

0

知识 (knowledge) · MoDELS · Performer · 样本 · 语言模型化 ·

2023 年 1 月 11 日

GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities

翻译：GPT作为知识工作者:对(AI)CPA能力进行零热评价

Jillian Bommarito,Michael Bommarito,Daniel Martin Katz,Jessica Katz

from arxiv, Source code and data available in online SI at https://github.com/mjbommar/gpt-as-knowledge-worker

The global economy is increasingly dependent on knowledge workers to meet the needs of public and private organizations. While there is no single definition of knowledge work, organizations and industry groups still attempt to measure individuals' capability to engage in it. The most comprehensive assessment of capability readiness for professional knowledge workers is the Uniform CPA Examination developed by the American Institute of Certified Public Accountants (AICPA). In this paper, we experimentally evaluate OpenAI's `text-davinci-003` and prior versions of GPT on both a sample Regulation (REG) exam and an assessment of over 200 multiple-choice questions based on the AICPA Blueprints for legal, financial, accounting, technology, and ethical tasks. First, we find that `text-davinci-003` achieves a correct rate of 14.4% on a sample REG exam section, significantly underperforming human capabilities on quantitative reasoning in zero-shot prompts. Second, `text-davinci-003` appears to be approaching human-level performance on the Remembering & Understanding and Application skill levels in the Exam absent calculation. For best prompt and parameters, the model answers 57.6% of questions correctly, significantly better than the 25% guessing rate, and its top two answers are correct 82.1% of the time, indicating strong non-entailment. Finally, we find that recent generations of GPT-3 demonstrate material improvements on this assessment, rising from 30% for `text-davinci-001` to 57% for `text-davinci-003`. These findings strongly suggest that large language models have the potential to transform the quality and efficiency of future knowledge work.

翻译：全球经济日益依赖知识工作者来满足公共和私营组织的需求。虽然对知识工作没有单一的定义,但各组织和行业团体仍然试图衡量个人参与知识工作的能力。对专业知识工作者能力准备情况的最全面评估是美国注册会计师协会(AICPA)开发的统一会计师考试。在本文中,我们实验性地评价OpenAI的“Text-davinci-003”和GPT的先前版本,对一项抽样条例(REG)考试以及对200多个基于AICPA法律、金融、会计、技术和道德任务蓝图的多种选择问题进行评估。首先,我们发现“text-davinci-003”在抽样REG考试部分实现了14.4%的正确率,在零点推论中大大低于定量推理的人力能力。第二,“text-da-davinci-003”似乎正在接近人类层面的绩效,在Examnational-ledgetal and Applemental legal disal disalth disalent 。为了最迅速和最准确的参数和最准确的答案是,在最新25.6%的答案中,我们最接近于25节正正确地发现,最新的30的答案。

0

相关内容

知识 (knowledge)

知识 (knowledge)

通过学习、实践或探索所获得的认识、判断或技能。

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

cGAS-STING信号通路在哮喘发病中的作用及其机制

国家自然科学基金

0+阅读 · 2014年12月31日

稀土MOF纳米荧光探针的设计合成及其生物应用

国家自然科学基金

0+阅读 · 2013年12月31日

外源性HCV RDRP对宿主细胞的表观遗传调控及在肝癌发病中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

NFκB信号通路调节巨噬细胞胆固醇平衡在尿毒症性动脉粥样硬化发病机制中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

CdTe/PbTe异质结二维电子气的电学特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

RTN3调控巨噬细胞自噬对动脉粥样硬化影响的研究

国家自然科学基金

0+阅读 · 2012年12月31日

可见及近红外宽光谱响应的高效固态量子点敏化太阳能电池

国家自然科学基金

0+阅读 · 2012年12月31日

新型稀土金属硼杂苯化合物化学

国家自然科学基金

0+阅读 · 2012年12月31日

鸡脾转移因子调节小肠黏液蛋白MUC2表达的"TLR-NOD"信号通路

国家自然科学基金

0+阅读 · 2012年12月31日

Notch信号通路负性调控哮喘小鼠气道杯状细胞MUC5AC的合成及其机制的研究

国家自然科学基金

0+阅读 · 2009年12月31日

Deconstructed Generation-Based Zero-Shot Model

Arxiv

0+阅读 · 2023年3月7日

Filter Pruning based on Information Capacity and Independence

Arxiv

0+阅读 · 2023年3月7日

Large Language Models as Zero-Shot Human Models for Human-Robot Interaction

Arxiv

1+阅读 · 2023年3月6日

Towards Zero-Shot Functional Compositionality of Language Models

Arxiv

0+阅读 · 2023年3月6日

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Arxiv

0+阅读 · 2023年3月6日

Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video Games

Arxiv

0+阅读 · 2023年3月2日

VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Arxiv

11+阅读 · 2022年12月26日

Updating Embeddings for Dynamic Knowledge Graphs

Arxiv

20+阅读 · 2021年9月22日

Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks

Arxiv

18+阅读 · 2021年6月17日

An Interpretable Reasoning Network for Multi-Relation Question Answering

Arxiv

13+阅读 · 2018年6月1日

VIP会员

文章信息

相关主题

知识 (knowledge)

语言模型化

相关VIP内容

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

俄乌战争启示：坦克战与不断演变的战斗形态

《大规模作战行动中与无人机集成的C5ISR系统》

《主观概率约束下寻找可行系统及其军事应用》69页

《美政府问责局：多种挑战影响地面战车任务出勤率》2025最新130页

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium6

中国图象图形学学会CSIG

2+阅读 · 2021年11月12日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium5

中国图象图形学学会CSIG

1+阅读 · 2021年11月11日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Deconstructed Generation-Based Zero-Shot Model

Arxiv

0+阅读 · 2023年3月7日

Filter Pruning based on Information Capacity and Independence

Arxiv

0+阅读 · 2023年3月7日

Large Language Models as Zero-Shot Human Models for Human-Robot Interaction

Arxiv

1+阅读 · 2023年3月6日

Towards Zero-Shot Functional Compositionality of Language Models

Arxiv

0+阅读 · 2023年3月6日

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Arxiv

0+阅读 · 2023年3月6日

Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video Games

Arxiv

0+阅读 · 2023年3月2日

VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Arxiv

11+阅读 · 2022年12月26日

Updating Embeddings for Dynamic Knowledge Graphs

Arxiv

20+阅读 · 2021年9月22日

Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks

Arxiv

18+阅读 · 2021年6月17日

An Interpretable Reasoning Network for Multi-Relation Question Answering

Arxiv

13+阅读 · 2018年6月1日

相关基金

cGAS-STING信号通路在哮喘发病中的作用及其机制

国家自然科学基金

0+阅读 · 2014年12月31日

稀土MOF纳米荧光探针的设计合成及其生物应用

国家自然科学基金

0+阅读 · 2013年12月31日

外源性HCV RDRP对宿主细胞的表观遗传调控及在肝癌发病中的作用

国家自然科学基金

0+阅读 · 2012年12月31日

NFκB信号通路调节巨噬细胞胆固醇平衡在尿毒症性动脉粥样硬化发病机制中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

CdTe/PbTe异质结二维电子气的电学特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

RTN3调控巨噬细胞自噬对动脉粥样硬化影响的研究

国家自然科学基金

0+阅读 · 2012年12月31日

可见及近红外宽光谱响应的高效固态量子点敏化太阳能电池

国家自然科学基金

0+阅读 · 2012年12月31日

新型稀土金属硼杂苯化合物化学

国家自然科学基金

0+阅读 · 2012年12月31日

鸡脾转移因子调节小肠黏液蛋白MUC2表达的"TLR-NOD"信号通路

国家自然科学基金

0+阅读 · 2012年12月31日

Notch信号通路负性调控哮喘小鼠气道杯状细胞MUC5AC的合成及其机制的研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员