面向高能物理应用的可扩展分布式计算系统评估的代理建模 (Surrogate Modeling for Scalable Evaluation of Distributed Computing Systems for HEP Applications)

from arxiv, Included in EPJ Web of Conferences Volume 337 (2025). 27th International Conference on Computing in High Energy and Nuclear Physics (CHEP 2024)

The Worldwide LHC Computing Grid (WLCG) provides the robust computing infrastructure essential for the LHC experiments by integrating global computing resources into a cohesive entity. Simulations of different compute models present a feasible approach for evaluating future adaptations that are able to cope with future increased demands. However, running these simulations incurs a trade-off between accuracy and scalability. For example, while the simulator DCSim can provide accurate results, it falls short on scaling with the size of the simulated platform. Using Generative Machine Learning as a surrogate presents a candidate for overcoming this challenge. In this work, we evaluate the usage of three different Machine Learning models for the simulation of distributed computing systems and assess their ability to generalize to unseen situations. We show that those models can predict central observables derived from execution traces of compute jobs with approximate accuracy but with orders of magnitude faster execution times. Furthermore, we identify potentials for improving the predictions towards better accuracy and generalizability.

翻译：全球大型强子对撞机计算网格（WLCG）通过整合全球计算资源形成一个统一整体，为大型强子对撞机实验提供了至关重要的稳健计算基础设施。对不同计算模型进行仿真是评估未来适应性调整的一种可行方法，使其能够应对日益增长的需求。然而，运行这些仿真需要在准确性与可扩展性之间进行权衡。例如，虽然仿真器DCSim能够提供精确结果，但在模拟平台规模扩展方面存在不足。利用生成式机器学习作为代理模型，为克服这一挑战提供了一种可能方案。在本工作中，我们评估了三种不同机器学习模型在分布式计算系统仿真中的应用，并考察了它们对未见场景的泛化能力。研究表明，这些模型能够以近似精度预测从计算作业执行轨迹中提取的核心观测指标，同时获得数量级级别的执行速度提升。此外，我们指出了通过改进预测精度和泛化能力来提升模型性能的潜在方向。

相关内容

MoDELS

关注 44

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日