My series of fully open, state-of-the-art small mixture-of-experts models.
🔄 In a Training Loop
aquilesfd
aquilesfd
AI & ML interests
research
Recent Activity
liked a model 3 minutes ago
ibm-granite/granite-4.2-30b liked a model about 24 hours ago
BananaMind/Overfitter-1.0 repliedto Banaxi-Tech's post 4 days ago
We're announcing our BananaMind 2.1 model series!
The models will include:
- BananaMind 2.1 Nano: 10M parameters with 60B tokens.
- BananaMind 2.1 Lite: 25M parameters with 40B tokens.
- BananaMind 2.1 Flash: 50M parameters with 55B tokens.
- BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens.
These model will use a multi tower architecture (like https://huggingface.co/BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
Follow us:
https://huggingface.co/BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
https://huggingface.co/bananamind-research-community