In complex auditory environments, humans can selectively attend to a target speaker while suppressing interfering sources. Electroencephalography (EEG)-based auditory attention decoding (AAD) aims to identify the speech source an individual is attending to, holding significant potential for applications in hearing aids and human-machine interaction. However, existing methods still exhibit notable limitations in modeling dynamic time-frequency features, extracting multi-scale representations, and computational efficiency.To address these challenges, we propose a dual-stream time-frequency convolutional network with additive attention. The temporal branch employs a temporal convolutional additive attention module to integrate temporal and channel attention, enhancing the modeling of key EEG time segments and feature channels. The spectral branch adopts a spatial-spectral feature extractor and a multi-scale feature enhancement module to improve sensitivity to and modeling of frequency variation patterns. The global features from both branches are ultimately fused for attention direction classification. Under a 1-second decision window, our model achieved accuracy rates of 96.5%, 86.3%, 73.1%, and 73.5% on the KUL, DTU, AVED (audio-only), and AVED (audio-visual) datasets, respectively. These results significantly outperform state-of-the-art methods and demonstrate the effectiveness and superiority of the proposed approach.