Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
34.8
TFLOPS
Wang Weiyi
PRO
kaupane
3
24
193
Follow
21world's profile picture
FlameF0X's profile picture
Fishtiks's profile picture
7 followers
·
46 following
Mtrya
AI & ML interests
None yet
Recent Activity
reacted
to
nwaughachukwuma
's
post
with 👍
2 days ago
Can a text-only model + a vision toolkit (mm-ctx) match a native vision model? We benchmarked 4 setups on 23 multimodal tasks (image, video, audio, PDF): • glm-5.2 (text-only) + mm-ctx: 88.4 • gemini-3.5-flash (vision): 83 • deepseek-v4-pro (text-only) + mm-ctx: 79.4 • qwen3.6-35b-a3b (vision): 44.3 The best text-only setup `glm-5.2 + mm` outperformed gemini-3.5-flash, the top vision model, by 5.4 points (6.5%). It was also: • 1.5x faster (100s vs 150s mean per task) • the only setup with zero timeouts (46/46 completed; gemini timed out 4x on bulk-image and long-video tasks) • the only setup stable across runs (88.5 / 88.4) • top on video (100.0), image (91.7), and PDF (90.0) tasks The trade-offs: the toolkit consumed 3.3x more tokens (4.25M vs 1.28M), and lost on audio (85.6 vs 71.3). On completed tasks alone the two are nearly identical (91.0 vs 88.4): the toolkit's edge is efficient extraction that keeps long media tasks inside the time budget. Full report: https://huggingface.co/blog/vlm-run/text-only-models-with-mm
liked
a model
8 days ago
moonshotai/Kimi-K3
liked
a Space
10 days ago
OpenMOSS-Team/MOSS-transcribe-diarize
View all activity
Organizations
None yet
kaupane
's models
6
Sort: Recently updated
kaupane/ArtFlow
0.7B
•
Updated
Feb 26
kaupane/DiT-Wikiart-Large
Text-to-Image
•
0.2B
•
Updated
Nov 1, 2025
•
14
•
1
kaupane/DiT-Wikiart-Small
Text-to-Image
•
22M
•
Updated
Nov 1, 2025
•
6
kaupane/DiT-Wikiart-Base
Text-to-Image
•
90.3M
•
Updated
Nov 1, 2025
•
2
kaupane/ChessFormer-RL
Reinforcement Learning
•
0.1B
•
Updated
Jun 28, 2025
•
4
•
1
kaupane/ChessFormer-SL
Reinforcement Learning
•
0.1B
•
Updated
Jun 28, 2025
•
7
•
1