The first <sigh> does sound a lot like a moan. OP linked to the timestamp so I missed it when it first played. I was also confused but on second playback I heard the first <sigh> and also thought wtf.
Well if possible you want an AI that understand the spirit of what you are asking for and will add all the missing stuff, instead of an AI that just want to hack its way to the result. Kinda what Fable brought to the table. For instance as a simple example I ask it to change the text that shows the email of the user by his name and Fable did all the code in case there is the family name missing etc. That this last part you want an AI to do. Helping you to build the system with you and not gaming what you ask for for reward.
You can do the same thing with celsius if you live in a country with a different temperature "calendar".
I'm used to the idea that 0C is what splits the year between cold season and warm season. Obviously people from California will see it differently, for example.
So where I live it is like this:
30 or more - hot
15..30 warm
0..15 chilly
-15..0 calendar winter
-30..-15 real winter
..-30 - freezing, kinds younger than 12 don't go to school
If only this infrastructure could handle all the traffic. I've tried using glm via z.ai - and it's a snail kind of slow.
And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
That it's slow doesn't mean it can't handle the traffic, just that this speed is the optimal tradeoff to them. They benefit from serving more tokens by exploiting parallelism across users at a lower number of tokens per second per user, instead of serving each individual user as quickly as possible. When there's a drop in traffic, they probably shut down GPUs rather than giving you higher speed.
It can very well be Sol, no? What stops them from using cheaper model for some requests during "rush" hours or simply use cheaper model for every Nth request.
>What stops them <..> simply use cheaper model for every Nth request.
That would trigger a full prefill (context recompute) every Nth request because cached tokens aren't interchangeable between models, and that would require way more compute than just staying on Astra.
To avoid full recompute, you could prefill a cheaper model's context incrementally by always feeding it Astra's outputs in the background (and vice versa), but then that would require 1.5-2 more VRAM for each session + the complexity of keeping them in sync.
If the rumors are true that Astra is a looped transformer, a more practical approach would be to dynamically adjust the loop count during peak hours.
>I go from woman to mother. And with it comes judgement and duties, and goes freedom and care-free living.
Agree on judgement, but duties and freedom and care-free living?
Sorry, both are gone for fathers too. At least those who care (talking as a father here).
Sure, a woman has even less freedom for a few years after a child is born. And a caring partner will try to at least ease this period for a woman.
But rest of a time? It's pretty 50/50 in my experience. My wife (also a backend dev like myself) has her career and her free time, in our family we share cleaning duties and all the cooking is on me (both because I'm better at it and it's less of a burden for me).
I'd like to eliminate judgement (and event more often self-judgement) she sometime faces, but for now I have no solution to present.
Bottom line: it's a fantasy to thing that a responsible father just carries on with whatever life he had before the kid was born and just occasionally fetches one from school to have a few word on the go.
Onomatopoeia? Sure it is there, and some fillers (or whatever you call those little sounds). But moans?
reply