Эволюция архитектур трансформеров: от самовнимания к эффективным Sparse-моделям
Loading...
Files
item.page.doi
Authors
item.page.other-contributor
item.page.advisors
item.page.editors
Date
Publisher
item.page.issn
item.page.isbn
item.page.depon
item.page.diss
Journal Title
Journal ISSN
Volume Title
item.page.title-alternative
Evolution of transformer architectures: from self-attention to efficient sparse models
Abstract
Статья посвящена исследованию новых архитектур искусственных нейронных сетей, направленных на улучшение классических трансформеров путем уменьшения вычислительных затрат и увеличения эффективности. Рассматриваются основные этапы эволюции трансформеров, начиная с введения механизма самовнимания и заканчивая современными Sparse-моделями. Особое внимание уделяется методам оптимизации, таким как параметрически эффективный fine-tuning (PEFT), адаптивные слои (adapters), а также использование внешних хранилищ знаний и mem-векторов для расширения долговременной памяти моделей. Статья завершается обсуждением текущих достижений, ограничений и перспективных направлений будущих исследований в данной области.
item.page.description-other
The article is devoted to the study of new architectures of artificial neural networks aimed at improving classical transformers by reducing computational costs and increasing efficiency. The main stages of the evolution of transformers are considered, starting with the introduction of the mechanism of self-attention and ending with modern sparse models. Particular attention is paid to optimization techniques such as parametrically efficient fine-tuning (PEFT), adaptive layers (adapters), as well as the use of external knowledge repositories and memory vectors to expand the long-term memory of models. The article concludes with a discussion of current achievements, limitations, and promising areas for future research in this area.
Description
Citation
Гуляев, В. О. Эволюция архитектур трансформеров: от самовнимания к эффективным Sparse-моделям = Evolution of transformer architectures: from self-attention to efficient sparse models / В. О. Гуляев, О. И. Захарова // Системный анализ и прикладная информатика. – 2026. – № 2. – С. 78-81.