CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset - 专知论文

会员服务 ·

0

CSS · 数据集 · 讲稿 · Parse · Performer ·

2023 年 5 月 25 日

CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical Dataset

翻译：暂无翻译

Hanchong Zhang,Jieyu Li,Lu Chen,Ruisheng Cao,Yunyan Zhang,Yu Huang,Yefeng Zheng,Kai Yu

The cross-domain text-to-SQL task aims to build a system that can parse user questions into SQL on complete unseen databases, and the single-domain text-to-SQL task evaluates the performance on identical databases. Both of these setups confront unavoidable difficulties in real-world applications. To this end, we introduce the cross-schema text-to-SQL task, where the databases of evaluation data are different from that in the training data but come from the same domain. Furthermore, we present CSS, a large-scale CrosS-Schema Chinese text-to-SQL dataset, to carry on corresponding studies. CSS originally consisted of 4,340 question/SQL pairs across 2 databases. In order to generalize models to different medical systems, we extend CSS and create 19 new databases along with 29,280 corresponding dataset examples. Moreover, CSS is also a large corpus for single-domain Chinese text-to-SQL studies. We present the data collection approach and a series of analyses of the data statistics. To show the potential and usefulness of CSS, benchmarking baselines have been conducted and reported. Our dataset is publicly available at \url{https://huggingface.co/datasets/zhanghanchong/css}.

翻译：暂无翻译

0

相关内容

CSS

层叠样式表（Cascading Style Sheet）是一种用来为结构化文档（如 HTML 文档或 XML 应用）添加样式（字体、间距和颜色等）的计算机语言。

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

专知会员服务

69+阅读 · 2020年1月2日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

LibRec 精选：推荐系统的常用数据集

LibRec 精选：推荐系统的常用数据集

LibRec智能推荐

17+阅读 · 2019年2月15日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

乳房外Paget病中错配修复基因功能异常与PI3K/AKT通路在侵袭中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

基于Notch信号通路及其表观遗传改变探讨养肺活血方防治肺纤维化的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

基于多性能退化参数的装备关键系统实时可靠性评估与预测方法研究

国家自然科学基金

1+阅读 · 2012年12月31日

基质中成纤维细胞在乳腺癌内分泌耐药中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

LNPEP基因与汉族人银屑病发病机制相关性研究

国家自然科学基金

0+阅读 · 2012年12月31日

UNITE: A Unified Benchmark for Text-to-SQL Evaluation

Arxiv

0+阅读 · 2023年7月14日

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Arxiv

0+阅读 · 2023年7月13日

Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events

Arxiv

0+阅读 · 2023年7月13日

A Survey for Biomedical Text Summarization: From Pre-trained to Large Language Models

Arxiv

0+阅读 · 2023年7月13日

Towards Expert-Level Medical Question Answering with Large Language Models

Arxiv

26+阅读 · 2023年5月16日

VIP会员

文章信息

相关主题

相关VIP内容

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

2019年自然语言处理NLP亮点总结，29页pdf，NLP Year in Review — 2019 NLP highlights for the year 2019.

专知会员服务

69+阅读 · 2020年1月2日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

前沿人工智能趋势报告（Frontier AI Trends Report）

【AAAI2026】善始则事半功倍：基于前缀优化的大语言模型推理强化学习

Andrej Karpathy：2025 年 LLM 年度回顾（2025 LLM Year in Review）

音退化问题：基于输入操控的鲁棒语音转换综述

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

LibRec 精选：推荐系统的常用数据集

LibRec 精选：推荐系统的常用数据集

LibRec智能推荐

17+阅读 · 2019年2月15日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

UNITE: A Unified Benchmark for Text-to-SQL Evaluation

Arxiv

0+阅读 · 2023年7月14日

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Arxiv

0+阅读 · 2023年7月13日

Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events

Arxiv

0+阅读 · 2023年7月13日

A Survey for Biomedical Text Summarization: From Pre-trained to Large Language Models

Arxiv

0+阅读 · 2023年7月13日

Towards Expert-Level Medical Question Answering with Large Language Models

Arxiv

26+阅读 · 2023年5月16日

相关基金

乳房外Paget病中错配修复基因功能异常与PI3K/AKT通路在侵袭中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

基于Notch信号通路及其表观遗传改变探讨养肺活血方防治肺纤维化的作用机制

国家自然科学基金

0+阅读 · 2012年12月31日

基于多性能退化参数的装备关键系统实时可靠性评估与预测方法研究

国家自然科学基金

1+阅读 · 2012年12月31日

基质中成纤维细胞在乳腺癌内分泌耐药中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

LNPEP基因与汉族人银屑病发病机制相关性研究

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员