Translating natural language into Bash Commands is an emerging research field that has gained attention in recent years. Most efforts have focused on producing more accurate translation models. To the best of our knowledge, only two datasets are available, with one based on the other. Both datasets involve scraping through known data sources (through platforms like stack overflow, crowdsourcing, etc.) and hiring experts to validate and correct either the English text or Bash Commands. This paper provides two contributions to research on synthesizing Bash Commands from scratch. First, we describe a state-of-the-art translation model used to generate Bash Commands from the corresponding English text. Second, we introduce a new NL2CMD dataset that is automatically generated, involves minimal human intervention, and is over six times larger than prior datasets. Since the generation pipeline does not rely on existing Bash Commands, the distribution and types of commands can be custom adjusted. We evaluate the performance of ChatGPT on this task and discuss the potential of using it as a data generator. Our empirical results show how the scale and diversity of our dataset can offer unique opportunities for semantic parsing researchers.
翻译:将自然语言转换成巴什指令是一个新兴的研究领域,近年来引起了人们的注意。 大部分努力都集中在制作更准确的翻译模型上。 根据我们的最佳知识,只有两个数据集可用,其中一个基于另一个。 两个数据集都涉及通过已知的数据源(通过堆叠溢、众包等平台)进行筛选,以及雇用专家验证和校正英文文本或巴什指令。本文为从头到尾合成巴什指令的研究提供了两项贡献。 首先,我们描述了用于从相应的英文文本中生成巴什指令的最先进的翻译模型。 其次,我们引入了一个新的NL2CMD数据集,该数据集自动生成,涉及最低限度的人类干预,比先前的数据集大六倍以上。由于生成管道不依赖现有的巴什指令,因此,命令的分布和类型可以自定调整。 我们评估了查普特在这项任务上的性能, 并讨论了使用它作为数据生成器的可能性。 我们的实证结果显示, 我们的数据集的规模和多样性能够为地震研究人员提供独特的机会。</s>