The exciting arrival of Mamba has created considerable attention within the machine learning community . This novel architecture, unlike conventional Transformers, offers a potential path to superior speed and diminished resource requirements. Departing from the quadratic scaling inherent in attention , Mamba leverages a structured method that seek