Hacker Newsnew | past | comments | ask | show | jobs | submit | anilgulecha's commentslogin

> Qwen2.5-Coder-3B

Basing it's findings of LLM as judge on this model, and then proceeding to ignore it. This article can be safely ignored as well.

LLM as judge in harness evals is the way to go, for any of your custom needs. Design the eval well.


I don't know where you pulled that from, its not in the article.


Follow the paper from the LLM as judge section.


I created my own in-browser terminal, so it can sit in chrome, accessible, and with some enhancements like animating the tab when the agent is doing some work. It's a layer on top of ttyd/tmux.

https://github.com/anilgulecha/ttydterm


> put safety above profit.

I think does not reflect the reality of how competitive capitalism works.


If "competitive capitalism" is how you want to frame it, then Zuckerberg should have been thrown out for essentially burning $83 billion dollars on a VR world that is never, ever going to pan out. That's going into a hole at over $400k per user. It's a disaster and any CEO of any functional company would be out on their ass, but Meta and Facebook are quite dysfunctional.


> So where does this leave frontend dev education?

I think this is more generally,

> So where does this leave education?

We need to figure this out anyway, because gpt galaxia and gpt cosmos will mean no skill is worth learning functionally.

The law of 42: Any debate regarding artificial intelligence and human skill will inevitably collapse into a discussion on the meaning of life, the universe, and everything.


Currently (re)learning my math fundamentals and it is indeed frustrating to see that Fable 5 has everything correct. It reads my scribbles and then tells me whether my answer is right or wrong.


Hard disagree.

The AI is a fantastic assistive tool, and in many ways it surpass my skillset as a software developer(I'm a skilled software developer, although it's not my job).

However, I do not believe I would be able to build what I'm building using LLMs, without myself having the right skills.


you forgot a 'yet' somewhere in the last sentence of your comment.


No, I didn't forget anything. Don't put words in my mouth.

I do not believe random Joe, even an intelligent one, will be able to build advanced software using LLMs in the near future.


Very good tool: I see the go support is not released. Will you be cutting a new release? Will take it for a spin. IMO you can specialize this as a read only AST/search tool, as you said, let models use regular edit tool.

Also some sample commands runs and output, would be useful in the README, for anyone passing by.


Yeah, I'll do a new release with a whole bunch of local fixes later today.


>ridiculous and abhorrent

Why such an adverse reaction? Nothing about what they said seems abhorrent! It's a common/usual positive sentiment, like "Understanding things give life meaning", etc.


Because getting "common language and culture" is not neutral when we reside in a world that basically this works with american culture being the global culture. Honestly imo this is a kind of perspective that only an american could have, and being ignorant about what it actually entails. You do not need everybody else to adopt a certain language and way of living to get to understand each other and have peace.

Moreover as other examples here given, geopolitical tensions very often appear between neighbours who have a very similar culture, so sharing cultural similarities does not seem to lead to much of peace in the world.


You seem to be projecting "this will impose western culture everywhere" on perhaps an LLM translater being used to ask if they "wanted to play board games at a cafe with me". Perhaps I didn't read all that into "common culture", a specific example was called out.

Cultures are not that weak, or done away with that easily.


Nobody is arguing that translator apps existing is a bad thing, that was not what was told. Obviously being able to communicate with each other is a good thing.

Not sure about nepal, but in some cultures/places (can just be different place within the same broadly culture) they associate eg card playing with gambling. I do not gamble but I play games with cards occasionally, and if I say "let's play a card game in a cafe" without meaning it for gambling it may still be interpreted as suggesting to go gamble by some, even though there is no literal translation issue per se (but there is misinterpretation). I dunno but maybe something similar is what happened.


Not the OP but I guess the statement implied cultural and lingual unification which implies erasure of other cultures and languages. We do have to work on being understood across the cultural and lingual division but the differences should be celebrated and welcome, even if they do cause occasional misunderstandings.


That's bragging rights correctly earned, i think! As marketing-y as this post is, definitely something to keep an eye on.


What incentive would they have when their customer also has the same option and will do it directly?


Do what? Customer searching for a physical book will do print-on-demand or what do you mean? Most customers will go with a cheap and convenient option


A free ebook download. What is cheaper and more convenient.


The premise of Nolan's position is Claude's language tic should not rub off on him, vs it's ok for Tolstoy's to do so.

I disagree. I think words are powerful, and when a word will be understood exactly as it's stated, then the value of the word is present. So while he might bemoan "load-bearing" (so would many of it's readers), that word in the right place will exactly communicate what he'd like. That enough reason to "earn it's keep" in prose.


it's not that load-bearing isn't doing the, well, load-bearing semantic work but it's both the lack of variability and lack of nuance hat makes it sound cheap. word choice is powerful - and in the same way, it can be nuanced

for eg. consider the synonyms for 'load-bearing'. if you want construction motifs: keystone, linchpin, cornerstone, foundational, structural. if you want earthy: bedrock, root. if you want ten-dollar words: axiomatic, sine qua non. if you just want to-the-point: vital, key, tenet, crucial

each of those words contain nuance. vital is used a lot in medicine (vital signs) so health, the heart, etc are evoked. pair it with specific usage like "this point is vital, the argument would die without it" and it rings different than "this point is load-bearing, the argument would be refuted without it"

you'd need to exist as a purely logical, probabilistic semantic parsing machine with no innate feeling or cultural attunement to prefer the lack of color and nuance. language as fully utilitarian with no artfulness is dreadful - if all I had to read all day was a self-important tech manual, I too would be frustrated and want for better craft in output prose


I didn't say I'd prefer it. I said that's the only word that presented itself to the author, it provided value, and there's no reason to feel bad about it.


The fact Anthropic, OpenAI and Google think they can fingerprint model output with no fidelity loss tells me otherwise.

As does caveman.


Why? We can also fingerprint human authors with high degrees of certainty if given large enough samples.

We've been doing so since at least before the media excitement over Robert Galbraith being outed via software analysis as J.K. Rowling.


> assuming you actually review what it outputs

You don't need to do this if you setup TDD. Over time, you develop confidence in the tests, and those ensure your code is solid. Linters/formatters ensure coding standards.

It's a shift, but once you make it, is when agentic development clicks.


TDD is fantastic. It's also insufficient on its own. Humans must verify code line-by-line - unless you're writing disposable code, which is fine! That has its place.

More and more, this is becoming an unreconcilable ideological difference in software development.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: