Yeah, tokens/s varies a lot based on workload. However, I’ve calibrated the estimator against public benchmarks, so it stays within a 30% error margin!
I'll definitely explore how to estimate vision models next!
We've been running a small finetuned VLM for OCR and yeah... cost is like 1/3rd of the usual vision APIs. Small models are kinda the whole product for us
That would be cool! Now there are only a few MCP servers enabled, but plan to make it exstensible. Do you have some apps you'd like to control via voice?
Love how this is just a chrome extension... Just made an API so that you could easily OCR everything with SOTA results (using finetuned VLM) at 1/3rd of the usual costs... Would absolutely love to chat and see if we can help out !