4比特精确度的例子:k比特推论法 (The case for 4-bit precision: k-bit Inference Scaling Laws) - 专知论文

会员服务 ·

0

查准率/准确率 · 缩放 · MoDELS · 推断 · 模型评估 ·

2022 年 12 月 19 日

The case for 4-bit precision: k-bit Inference Scaling Laws

翻译：4比特精确度的例子:k比特推论法

Tim Dettmers,Luke Zettlemoyer

Quantization methods reduce the number of bits required to represent each parameter in a model, trading accuracy for smaller memory footprints and inference latencies. However, the final model size depends on both the number of parameters of the original model and the rate of compression. For example, a 30B 8-bit model and a 60B 4-bit model have the same number of bits but may have very different zero-shot accuracies. In this work, we study this trade-off by developing inference scaling laws of zero-shot performance in Large Language Models (LLMs) to determine the bit-precision and model size that maximizes zero-shot performance. We run more than 35,000 zero-shot experiments with 16-bit inputs and k-bit parameters to examine which quantization methods improve scaling for 3 to 8-bit precision at scales of 19M to 66B parameters across the LLM families BLOOM, OPT, NeoX/Pythia, and GPT-2. We find that it is challenging to improve the bit-level scaling trade-off, with the only improvements being the use of a small block size -- splitting the parameters into small independently quantized blocks -- and the quantization data type being used (e.g., Int vs Float). Overall, our findings show that 4-bit precision is almost universally optimal for total model bits and zero-shot accuracy.

翻译：量化方法减少了在模型中代表每个参数所需的比特数数量,将精确度转换成较小的记忆足迹和推推延迟。然而,最终模型大小取决于原始模型参数的数量和压缩率。例如,30B 8比特模型和60B 4比特模型的比特数数量相同,但可能具有非常不同的零射偏差。在这项工作中,我们通过在大语言模型中制定零射性能的推论测量法来研究这一权衡,以确定使零弹性能最大化的比特精度和模型大小。我们用16比特投入和k比特参数进行超过35 000次零射试验,以检查在19M至66比特标准范围内,在LLOM家族BLOM、ALMO、NoX/Pythia和GPT-2之间,哪些四位化方法改进了3至8比特精确度。我们发现改进大语言模型零弹分级贸易的比特度和模型规模,只有利用小块精确度的改进,几乎是小块的精确度,而独立地使用了小块的平面图。

0

相关内容

查准率/准确率

查准率/准确率

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【2021新书】并行高性能计算，705页pdf，Parallel and High Performance Computing

【2021新书】并行高性能计算，705页pdf，Parallel and High Performance Computing

专知会员服务

106+阅读 · 2021年10月30日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

BERT到底如何work的？A Primer in BERTology: What we know about how BERT works

BERT到底如何work的？A Primer in BERTology: What we know about how BERT works

专知会员服务

50+阅读 · 2020年2月28日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

基于SiPM的高性能In-Beam TOF-PET的研究

国家自然科学基金

0+阅读 · 2014年12月31日

p53基因突变促进Wilms 肿瘤发展转移的小鼠动物模型研究

国家自然科学基金

0+阅读 · 2014年12月31日

失效物理模式下融合多源信息的空间滚动轴承服役可靠性研究

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

缺血性脑卒中:全脑氧代谢及微循环代偿机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

使用双能量CT建立监测贝伐单抗治疗非小细胞肺癌肿瘤内部变化标准的动物及临床研究

国家自然科学基金

0+阅读 · 2012年12月31日

细胞凋亡过程中半胱氨酸蛋白酶级联激活反应的分子成像

国家自然科学基金

0+阅读 · 2012年12月31日

肝癌细胞膜蛋白cytokeratin-1用于肝癌在体分子显像和靶向治疗的相关研究

国家自然科学基金

0+阅读 · 2011年12月31日

A Simplistic Model of Neural Scaling Laws: Multiperiodic Santa Fe Processes

Arxiv

0+阅读 · 2023年2月17日

Fast and Robust Non-Rigid Registration Using Accelerated Majorization-Minimization

Arxiv

0+阅读 · 2023年2月16日

Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators

Arxiv

0+阅读 · 2023年2月16日

Bolstering Stochastic Gradient Descent with Model Building

Arxiv

0+阅读 · 2023年2月15日

Video Probabilistic Diffusion Models in Projected Latent Space

Arxiv

2+阅读 · 2023年2月15日

Direct multiple shooting and direct collocation show similar performances in biomechanical predictive simulations

Arxiv

0+阅读 · 2023年2月15日

Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models

Arxiv

0+阅读 · 2023年2月14日

Score-based Diffusion Models in Function Space

Arxiv

0+阅读 · 2023年2月14日

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Arxiv

1+阅读 · 2023年2月14日

TinyBERT: Distilling BERT for Natural Language Understanding

TinyBERT: Distilling BERT for Natural Language Understanding

Arxiv

11+阅读 · 2019年9月23日

VIP会员

文章信息

相关主题

查准率/准确率

相关VIP内容

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

【2021新书】并行高性能计算，705页pdf，Parallel and High Performance Computing

【2021新书】并行高性能计算，705页pdf，Parallel and High Performance Computing

专知会员服务

106+阅读 · 2021年10月30日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

BERT到底如何work的？A Primer in BERTology: What we know about how BERT works

BERT到底如何work的？A Primer in BERTology: What we know about how BERT works

专知会员服务

50+阅读 · 2020年2月28日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《陆军战斗操练中的关键事件诊断》

《自适应训练辅助概念及其在空战管理员加速训练中的应用导论》最新126页

军事通信市场七大趋势概述

《抗干扰无人机蜂群行为的遗传算法方法》

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

ACM TOMM Call for Papers

ACM TOMM Call for Papers

CCF多媒体专委会

2+阅读 · 2022年3月23日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

相关论文

A Simplistic Model of Neural Scaling Laws: Multiperiodic Santa Fe Processes

Arxiv

0+阅读 · 2023年2月17日

Fast and Robust Non-Rigid Registration Using Accelerated Majorization-Minimization

Arxiv

0+阅读 · 2023年2月16日

Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators

Arxiv

0+阅读 · 2023年2月16日

Bolstering Stochastic Gradient Descent with Model Building

Arxiv

0+阅读 · 2023年2月15日

Video Probabilistic Diffusion Models in Projected Latent Space

Arxiv

2+阅读 · 2023年2月15日

Direct multiple shooting and direct collocation show similar performances in biomechanical predictive simulations

Arxiv

0+阅读 · 2023年2月15日

Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models

Arxiv

0+阅读 · 2023年2月14日

Score-based Diffusion Models in Function Space

Arxiv

0+阅读 · 2023年2月14日

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Arxiv

1+阅读 · 2023年2月14日

TinyBERT: Distilling BERT for Natural Language Understanding

TinyBERT: Distilling BERT for Natural Language Understanding

Arxiv

11+阅读 · 2019年9月23日

相关基金

基于SiPM的高性能In-Beam TOF-PET的研究

国家自然科学基金

0+阅读 · 2014年12月31日

p53基因突变促进Wilms 肿瘤发展转移的小鼠动物模型研究

国家自然科学基金

0+阅读 · 2014年12月31日

失效物理模式下融合多源信息的空间滚动轴承服役可靠性研究

国家自然科学基金

0+阅读 · 2013年12月31日

半导体衬底上FeSe薄膜的外延生长及界面超导

国家自然科学基金

0+阅读 · 2013年12月31日

Schrodinger-Poisson方程的若干问题研究

国家自然科学基金

1+阅读 · 2012年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

缺血性脑卒中:全脑氧代谢及微循环代偿机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

使用双能量CT建立监测贝伐单抗治疗非小细胞肺癌肿瘤内部变化标准的动物及临床研究

国家自然科学基金

0+阅读 · 2012年12月31日

细胞凋亡过程中半胱氨酸蛋白酶级联激活反应的分子成像

国家自然科学基金

0+阅读 · 2012年12月31日

肝癌细胞膜蛋白cytokeratin-1用于肝癌在体分子显像和靶向治疗的相关研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员