Inference Providers
Active filters: trl
santiviquez/reward_modeling_anthropic_hh
Text Classification
• 0.3B • Updated • 60
• 7
raghu1155/DeepSeek-R1-Codegeneration-COT
Text Generation
• Updated • 1.37k
• 10
blai88/reward_modeling_anthropic_hh
0.3B • Updated • 36
• 3
shailja/lora_codellm_34b_verilog_model
mradermacher/DeepSeek-R1-Codegeneration-COT-GGUF
8B • Updated • 1.03k
• 10
mradermacher/gpt-4o-distil-Llama-3.1-8B-Instruct-PaperWitch-heresy-GGUF
8B • Updated • 819
• 10
SpeculativeDecoding/doctorboom-qwen2.5-coder-7b-lora
wxzhang/dpo-selective-redteaming
Text Generation
• 7B • Updated • 105
• 5
SiMajid/value_reward_modeling
Text Classification
• 0.3B • Updated • 14
• 3
trentmkelly/gpt-4o-distil-Llama-3.1-8B-Instruct
Text Generation
• Updated • 25
• 3
CEIA-RL/energy-exp1-dpo-offline
Text Generation
• 4B • Updated • 120
• 6
Text Generation
• 8B • Updated • 26
• 2
ram-lexsi/agenttune-testrun-tree-of-thoughts
sthanika-ai/gemma3-12b-kcc-advisory
Text Generation
• Updated • 15
• 3
ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-0.8B
Text Generation
• 0.8B • Updated • 4.35k
• 14
ConicCat/Gemma4-Writer-26BA4B
Image-Text-to-Text
• 26B • Updated • 72
• 3
773bw-h/my_first_reward_modeling
Text Classification
• 0.3B • Updated • 18
• 3
APaul1/Llama-3-8B-sft-lora-ultrachat
Updated • 12
• 1
HuggingFaceTB/SmolLM-135M-Instruct
Text Generation
• 0.1B • Updated • 28.2k
• 146
HuggingFaceTB/smollm-135M-instruct-v0.2-Q8_0-GGUF
0.1B • Updated • 3.19k
• 8
mlabonne/TwinLlama-3.1-8B-DPO
Text Generation
• 8B • Updated • 127
• 25
omersaidd/Prompt-Enhace-T5-base
0.2B • Updated • 21
• 3
dhirajlochib/llama-3.2-unsensored-3b
Updated • 12
abhinavasr/Hindu-Veda-Llama-3.2-3B-Instruct
3B • Updated • 3
linkred/stock_prediction_v5
Updated • 14
• 13
linkred/stock_prediction_v6
Updated • 15
• 12
linkred/stock_prediction_v7
Updated • 14
• 12
linkred/stock_prediction_v3_mini
Updated • 25
• 13
Masa1028/gpt2-instruction-tuning-alpaca
Text Generation
• 0.1B • Updated • 26
• 1
Locutusque/Thespis-Llama-3.1-8B
Text Generation
• 8B • Updated • 51
• • 16