MMTM: 数学字问题多学制多德变换器 (MMTM: Multi-Tasking Multi-Decoder Transformer for Math Word Problems)

Recently, quite a few novel neural architectures were derived to solve math word problems by predicting expression trees. These architectures varied from seq2seq models, including encoders leveraging graph relationships combined with tree decoders. These models achieve good performance on various MWPs datasets but perform poorly when applied to an adversarial challenge dataset, SVAMP. We present a novel model MMTM that leverages multi-tasking and multi-decoder during pre-training. It creates variant tasks by deriving labels using pre-order, in-order and post-order traversal of expression trees, and uses task-specific decoders in a multi-tasking framework. We leverage transformer architectures with lower dimensionality and initialize weights from RoBERTa model. MMTM model achieves better mathematical reasoning ability and generalisability, which we demonstrate by outperforming the best state of the art baseline models from Seq2Seq, GTS, and Graph2Tree with a relative improvement of 19.4% on an adversarial challenge dataset SVAMP.

翻译：最近,为通过预测表达式树来解决数学字问题,产生了一些新颖的神经结构。这些结构与随后的2Seq 模型不同,包括以图形为杠杆的编码器与树的解码器。这些模型在各种 MWP 数据集上表现良好,但在应用对抗性挑战数据集 SVAMP 时表现不佳。我们展示了一个新的MMTM 模型,该模型在培训前运用多任务和多解码功能,在培训前运用多种任务和多解码工具。它通过利用表达式树的顺序前、顺序和顺序后跨行来生成标签来创造不同的任务,并在多任务框架中使用特定的任务解码器。我们利用了低维度的变压器结构,并初始化了RoBERTa 模型的重量。 MMTM 模型实现了更好的数学推理能力和可概括性。我们通过在Seq2Seqeq、GTS和Sq2Treet2Te的艺术基线模型的最佳状态,在对抗性数据设置 SVAMP上相对改进了19.4%来证明我们表现最好的基准模型。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

53+阅读 · 2021年1月20日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日