Open Access
Issue
EPJ Web Conf.
Volume 380, 2026
International Conference on Information Systems and Communication Technologies (ICISCT’25)
Article Number 02002
Number of page(s) 13
Section Artificial Intelligence, Advanced Control Systems, and Energy Management
DOI https://doi.org/10.1051/epjconf/202638002002
Published online 03 August 2026
  1. J. Kaplan, J. McCandlish, T. Henighan, T.B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, D. Amodei, Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361 (2020). https://doi.org/10.48550/arXiv.2001.08361 [Google Scholar]
  2. deepseek ai, deepseek-v3 technical report. arxiv preprint arxiv:2412.19437 (2024). https://doi.org/10.48550/arXiv.2412.19437 [Google Scholar]
  3. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, I. Polosukhin, Attention Is All You Need. Advances in Neural Information Processing Systems 30 (2017). https://doi.org/10.48550/arXiv.1706.03762 [Google Scholar]
  4. A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, Improving Language Understanding by Generative Pre-Training. OpenAI Technical Report (2018). https://doi.org/10.21428/d587d40a.5d9e33d2 [Google Scholar]
  5. A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language Models are Unsupervised Multitask Learners. OpenAI Technical Report (2019). https://doi.org/10.21428/d587d40a.5d9e33d2 [Google Scholar]
  6. T. Brown et al., Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 33, 1877–1901 (2020). https://doi.org/10.48550/arXiv.2005.14165 [Google Scholar]
  7. D. Patterson et al., Carbon Emissions and Large Neural Network Training. arXiv preprint arXiv:2104.10350 (2021). https://doi.org/10.48550/arXiv.2104.10350 [Google Scholar]
  8. OpenAI, GPT-4 Technical Report. arXiv preprint arXiv:2303.08774 (2023). https://doi.org/10.48550/arXiv.2303.08774 [Google Scholar]
  9. J. Su, Y. Cao, D. Murthaza, S. Wen, Y. Chang, B. Wang, RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv preprint arXiv:2104.09864 (2021). https://doi.org/10.48550/arXiv.2104.09864 [Google Scholar]
  10. F. Gloeckle, B. Youbi Idrissi, B. Rozière, D. Lopez-Paz, G. Synnaeve, Better &Faster Large Language Models via Multi-token Prediction. arXiv preprint arXiv:2404.19737 (2024). https://doi.org/10.48550/arXiv.2404.19737 [Google Scholar]
  11. D. Dai et al., Deepseekmoe: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. CoRR abs/2401.06066 (2024). https://doi.org/10.48550/arXiv.2401.06066 [Google Scholar]
  12. DeepSeek AI, DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Ex-perts Language Model. CoRR abs/2405.04434 (2024). https://doi.org/10.48550/arXiv.2405.04434 [Google Scholar]
  13. N. Shazeer et al., Outrageously Large Neural Networks: The Sparsely-Gated Mix-ture-of-Experts Layer. Proceedings of the International Conference on Learning Representations (ICLR) (2017). https://doi.org/10.48550/arXiv.1701.06538 [Google Scholar]
  14. W. Fedus, B. Zoph, N. Shazeer, Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. arXiv preprint arXiv:2101.03961 (2021). https://doi.org/10.48550/arXiv.2101.03961 [Google Scholar]
  15. L. Wang, H. Gao, C. Zhao, X. Sun, D. Dai, Auxiliary Loss-Free Load Balancing Strategy for Mixture-of-Experts. arXiv preprint arXiv:2408.15664 (2024). https://doi.org/10.48550/arXiv.2408.15664 [Google Scholar]

Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.

Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.

Initial download of the metrics may take a while.