PaperScope
LIVE · 2026-09-30 05:40 UTC

Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms

Ene Meco, Emadeldeen Hamdan, A. Enis Cetin

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36698 v1
Category
Submitted
2026-09-29

Abstract

To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.

arXiv abs page · PDF · same-day batch