评估ChatGPT和GPT-4的逻辑推理能力 (Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4) - 专知论文

会员服务 ·

0

逻辑推理 · GPT-4 · 基准测试 · ChatGPT · 数据集 ·

2023 年 4 月 20 日

Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

翻译：评估ChatGPT和GPT-4的逻辑推理能力

Hanmeng Liu,Ruoxi Ning,Zhiyang Teng,Jian Liu,Qiji Zhou,Yue Zhang

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "advanced" at reasoning tasks, we are eager to learn the GPT-4 performance on various logical reasoning tasks. This report analyses multiple logical reasoning datasets, with popular benchmarks like LogiQA and ReClor, and newly-released datasets like AR-LSAT. We test the multi-choice reading comprehension and natural language inference tasks with benchmarks requiring logical reasoning. We further construct a logical reasoning out-of-distribution dataset to investigate the robustness of ChatGPT and GPT-4. We also make a performance comparison between ChatGPT and GPT-4. Experiment results show that ChatGPT performs significantly better than the RoBERTa fine-tuning method on most logical reasoning benchmarks. With early access to the GPT-4 API we are able to conduct intense experiments on the GPT-4 model. The results show GPT-4 yields even higher performance on most logical reasoning datasets. Among benchmarks, ChatGPT and GPT-4 do relatively well on well-known datasets like LogiQA and ReClor. However, the performance drops significantly when handling newly released and out-of-distribution datasets. Logical reasoning remains challenging for ChatGPT and GPT-4, especially on out-of-distribution and natural language inference datasets. We release the prompt-style logical reasoning datasets as a benchmark suite and name it LogiEval.

翻译：利用逻辑推理能力是一个广泛的自然语言理解任务。随着Generative Pretrained Transformer 4 (GPT-4)的发布，在推理任务上被标记为“高级”，我们渴望了解GPT-4在各种逻辑推理任务上的表现。该报告分析了多个逻辑推理数据集，包括流行的LogiQA和ReClor基准测试，以及新发布的AR-LSAT数据集。我们使用需要逻辑推理的基准测试来测试多项选择阅读理解和自然语言推理任务。此外，我们还构建了一个逻辑推理超出分布数据集，以研究ChatGPT和GPT-4的稳健性。我们还对比了ChatGPT和GPT-4的性能。实验结果表明，ChatGPT在大多数逻辑推理基准测试中的表现明显优于RoBERTa精调方法。通过提前访问GPT-4 API，我们能够对GPT-4模型进行深入的实验。结果表明，GPT-4在大多数逻辑推理数据集上的性能表现更好。在基准测试中，ChatGPT和GPT-4在像LogiQA和ReClor这样的知名数据集上表现得相对较好。但是，在处理新发布的和超出分布数据集时，性能显著下降。逻辑推理对于ChatGPT和GPT-4仍然具有挑战性，特别是在超出分布和自然语言推理数据集上。我们发布了提示式逻辑推理数据集作为基准测试套件，并将其命名为LogiEval。

2

相关内容

逻辑推理

ChatGPT和GPT-4的逻辑推理如何？浙大等最新《ChatGPT和GPT-4逻辑推理能力全面评测》论文解答，常规优异新数据差

ChatGPT和GPT-4的逻辑推理如何？浙大等最新《ChatGPT和GPT-4逻辑推理能力全面评测》论文解答，常规优异新数据差

专知会员服务

65+阅读 · 2023年4月19日

揭秘ChatGPT情感对话能力

揭秘ChatGPT情感对话能力

专知会员服务

59+阅读 · 2023年4月9日

GPT-4在医学上能力如何？微软OpenAI《GPT-4在医疗难题上的能力》论文

GPT-4在医学上能力如何？微软OpenAI《GPT-4在医疗难题上的能力》论文

专知会员服务

115+阅读 · 2023年3月24日

【ACL2022-华盛顿大学】生成知识促进常识推理，Generated Knowledge Prompting for Commonsense Reasoning

【ACL2022-华盛顿大学】生成知识促进常识推理，Generated Knowledge Prompting for Commonsense Reasoning

专知会员服务

26+阅读 · 2022年3月1日

【USC2021】常识推理，47页ppt，Commonsense Reasoning in the Wild

专知会员服务

33+阅读 · 2021年10月9日

【CIKM2020】神经逻辑推理，Neural Logic Reasoning

【CIKM2020】神经逻辑推理，Neural Logic Reasoning

专知会员服务

51+阅读 · 2020年8月25日

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

专知会员服务

65+阅读 · 2020年5月12日

【ACL2020-浙大-微软】多轮对话推理数据集，MuTual: A Dataset for Multi-Turn Dialogue Reasoning

【ACL2020-浙大-微软】多轮对话推理数据集，MuTual: A Dataset for Multi-Turn Dialogue Reasoning

专知会员服务

37+阅读 · 2020年4月10日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

揭秘ChatGPT情感对话能力

揭秘ChatGPT情感对话能力

专知

16+阅读 · 2023年4月9日

赛尔笔记 | 逻辑推理阅读理解任务及方法

赛尔笔记 | 逻辑推理阅读理解任务及方法

哈工大SCIR

1+阅读 · 2022年6月7日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

赛尔原创@ACL 2022 | e-CARE: 可解释的因果推理数据集

赛尔原创@ACL 2022 | e-CARE: 可解释的因果推理数据集

哈工大SCIR

1+阅读 · 2022年5月12日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

自然语言处理常识推理综述论文，60页pdf

自然语言处理常识推理综述论文，60页pdf

专知

73+阅读 · 2019年4月4日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

岩藻糖基转移酶对非天然供体底物选择性的分子改造

国家自然科学基金

0+阅读 · 2014年12月31日

ME1介导的代谢重组在基底样乳腺癌的作用和机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

扩展的线性时段不变式的模型检验

国家自然科学基金

1+阅读 · 2014年12月31日

陆地碳数据同化中的模型“异参同效”问题研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于深层神经网络的多模态快速稀疏表征器

国家自然科学基金

3+阅读 · 2014年12月31日

20世纪50年代以来青藏高原气温变化的不确定性定量评估

国家自然科学基金

1+阅读 · 2013年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

聚电解质的表征

国家自然科学基金

0+阅读 · 2011年12月31日

铁电材料断裂的相场模拟和压电模式原子力显微镜表征

国家自然科学基金

0+阅读 · 2009年12月31日

三维模型语义分析与检索研究

国家自然科学基金

2+阅读 · 2008年12月31日

Deductive Verification of Chain-of-Thought Reasoning

Deductive Verification of Chain-of-Thought Reasoning

Arxiv

1+阅读 · 2023年6月6日

A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Arxiv

1+阅读 · 2023年6月5日

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Arxiv

0+阅读 · 2023年6月5日

Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning

Arxiv

0+阅读 · 2023年6月4日

Evaluating Language Models for Mathematics through Interactions

Arxiv

0+阅读 · 2023年6月2日

An Evaluation of Log Parsing with ChatGPT

Arxiv

1+阅读 · 2023年6月2日

Examining the Causal Effect of First Names on Language Models: The Case of Social Commonsense Reasoning

Arxiv

0+阅读 · 2023年6月1日

True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-4

Arxiv

0+阅读 · 2023年6月1日

LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities

Arxiv

21+阅读 · 2023年5月22日

A Survey of Knowledge Enhanced Pre-trained Models

Arxiv

28+阅读 · 2021年10月1日

VIP会员

文章信息

相关主题

相关VIP内容

ChatGPT和GPT-4的逻辑推理如何？浙大等最新《ChatGPT和GPT-4逻辑推理能力全面评测》论文解答，常规优异新数据差

ChatGPT和GPT-4的逻辑推理如何？浙大等最新《ChatGPT和GPT-4逻辑推理能力全面评测》论文解答，常规优异新数据差

专知会员服务

65+阅读 · 2023年4月19日

揭秘ChatGPT情感对话能力

揭秘ChatGPT情感对话能力

专知会员服务

59+阅读 · 2023年4月9日

GPT-4在医学上能力如何？微软OpenAI《GPT-4在医疗难题上的能力》论文

GPT-4在医学上能力如何？微软OpenAI《GPT-4在医疗难题上的能力》论文

专知会员服务

115+阅读 · 2023年3月24日

【ACL2022-华盛顿大学】生成知识促进常识推理，Generated Knowledge Prompting for Commonsense Reasoning

【ACL2022-华盛顿大学】生成知识促进常识推理，Generated Knowledge Prompting for Commonsense Reasoning

专知会员服务

26+阅读 · 2022年3月1日

【USC2021】常识推理，47页ppt，Commonsense Reasoning in the Wild

专知会员服务

33+阅读 · 2021年10月9日

【CIKM2020】神经逻辑推理，Neural Logic Reasoning

【CIKM2020】神经逻辑推理，Neural Logic Reasoning

专知会员服务

51+阅读 · 2020年8月25日

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

专知会员服务

65+阅读 · 2020年5月12日

【ACL2020-浙大-微软】多轮对话推理数据集，MuTual: A Dataset for Multi-Turn Dialogue Reasoning

【ACL2020-浙大-微软】多轮对话推理数据集，MuTual: A Dataset for Multi-Turn Dialogue Reasoning

专知会员服务

37+阅读 · 2020年4月10日

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

【微软雷德蒙研究院】小样本自然语言生成，Few-shot Natural Language Generation for Task-Oriented Dialog

专知会员服务

33+阅读 · 2020年2月29日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

热门VIP内容

开通专知VIP会员享更多权益服务

《战区安全决策课程体系》最新244页

《"无人机航母"原型平台》

任务规划与地形分析：现代复杂环境作战导航体系

《攻击场景描述形式化模型研究》

相关资讯

揭秘ChatGPT情感对话能力

揭秘ChatGPT情感对话能力

专知

16+阅读 · 2023年4月9日

赛尔笔记 | 逻辑推理阅读理解任务及方法

赛尔笔记 | 逻辑推理阅读理解任务及方法

哈工大SCIR

1+阅读 · 2022年6月7日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

赛尔原创@ACL 2022 | e-CARE: 可解释的因果推理数据集

赛尔原创@ACL 2022 | e-CARE: 可解释的因果推理数据集

哈工大SCIR

1+阅读 · 2022年5月12日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

自然语言处理常识推理综述论文，60页pdf

自然语言处理常识推理综述论文，60页pdf

专知

73+阅读 · 2019年4月4日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Deductive Verification of Chain-of-Thought Reasoning

Deductive Verification of Chain-of-Thought Reasoning

Arxiv

1+阅读 · 2023年6月6日

A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Arxiv

1+阅读 · 2023年6月5日

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Arxiv

0+阅读 · 2023年6月5日

Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning

Arxiv

0+阅读 · 2023年6月4日

Evaluating Language Models for Mathematics through Interactions

Arxiv

0+阅读 · 2023年6月2日

An Evaluation of Log Parsing with ChatGPT

Arxiv

1+阅读 · 2023年6月2日

Examining the Causal Effect of First Names on Language Models: The Case of Social Commonsense Reasoning

Arxiv

0+阅读 · 2023年6月1日

True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-4

Arxiv

0+阅读 · 2023年6月1日

LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities

Arxiv

21+阅读 · 2023年5月22日

A Survey of Knowledge Enhanced Pre-trained Models

Arxiv

28+阅读 · 2021年10月1日

相关基金

岩藻糖基转移酶对非天然供体底物选择性的分子改造

国家自然科学基金

0+阅读 · 2014年12月31日

ME1介导的代谢重组在基底样乳腺癌的作用和机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

扩展的线性时段不变式的模型检验

国家自然科学基金

1+阅读 · 2014年12月31日

陆地碳数据同化中的模型“异参同效”问题研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于深层神经网络的多模态快速稀疏表征器

国家自然科学基金

3+阅读 · 2014年12月31日

20世纪50年代以来青藏高原气温变化的不确定性定量评估

国家自然科学基金

1+阅读 · 2013年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

聚电解质的表征

国家自然科学基金

0+阅读 · 2011年12月31日

铁电材料断裂的相场模拟和压电模式原子力显微镜表征

国家自然科学基金

0+阅读 · 2009年12月31日

三维模型语义分析与检索研究

国家自然科学基金

2+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员