Soundata: 一个用于复制使用音频数据集的 Python 库 (Soundata: A Python library for reproducible use of audio datasets)

Soundata is a Python library for loading and working with audio datasets in a standardized way, removing the need for writing custom loaders in every project, and improving reproducibility by providing tools to validate data against a canonical version. It speeds up research pipelines by allowing users to quickly download a dataset, load it into memory in a standardized and reproducible way, validate that the dataset is complete and correct, and more. Soundata is based and inspired on mirdata and design to complement mirdata by working with environmental sound, bioacoustic and speech datasets, among others. Soundata was created to be easy to use, easy to contribute to, and to increase reproducibility and standardize usage of sound datasets in a flexible way.

翻译：Soundata是一个Python图书馆,用于以标准化的方式装载和操作音频数据集,消除每个项目对自定义装货机的需要,并通过提供工具,对照卡通版本验证数据,改善再复制性。它加速研究管道,使用户能够快速下载数据集,以标准化和可复制的方式将数据集装入记忆,验证数据集是完整和正确的,而且更多。Soundata以虚拟数据和设计为基础,并受到启发,通过与无害环境、生物声学和语音数据集等合作,补充虚拟数据。创造Soundata是为了便于使用、容易促进和更加灵活地复制和统一使用声音数据集。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【干货书】Python深度学习第二版，Deep Learning with Python, Second Edition

专知会员服务

171+阅读 · 2020年5月9日