The thinking people who would find this interesting and read this are probably more than capable of understanding this and critical enough to expect that. Conversation over the title is distraction of what's important. Just stick to keeping original source title and let people vote and down vote if they don't like it. That's what votes are for.
It's regularly the case that they are simply too long to fit, so editorializing is necessary. Either cutting them slightly short, removing descriptors, excess verbs etc, is better than chopping a word in half.
Unfortunately, advertisers are getting smarter and using bots to praise their own products on Reddit. Thanks to training on genuine comments, some models are very good at sounding like a human commenter, and can easily generate a comment history with diverse interests to appear human, making them basically undetectable. So it seems like this method will, at some point, not identify the company with the best knife, but the one with the most ad spend on bot comments.
I think this is the way. An LLM is an expensive general purpose tool and for repeatable tasks, after it's clarified the process flow, it builds cheaper special purpose tools for each step
100%. We now have super general tools that reduce the cost to build other specific tools. My favorite thing with LLMs has been building a ton of little utilities for work and personal that I could've built before, but never had the time at work or the want to spend time on in my personal time.
The other day I wanted to gather Reddit comments about a solar panel vendor. Claude doesn't have access to I had Gemini do some "deep research". When I fed the verbose report back to Claude it basically said it was a bunch of "hallucinated bullshit".
I've never had a Gemini Deep Research report that didn't sound like a load of pseudo-intellectual BS. It always starts with a long grandiose preamble and then sounds way too academic, almost like a caricature of academia.
I found it much more useful to go to a knife shop and handle a whole bunch of knives for myself. They’re all pretty similar besides material, so not much signal you’re going to be able to glean from people arguing on reddit.
Most of the attributes don’t matter. Most people would be much better off with a $50 Victorinox that they kept sharp and a wood cutting board they maintained than upgrading the knife. If you are using it all day there are definitely looking things from a comfort perspective but for most homes, does not matter.
Isn't it just learning to map specific words, from the "knife world", to the correct class? If so, a simple dictionary would fit.
What I think is a better way to validate is to split train/validation by words used presented in NER classes (like, it should be able to find new brands never seen before). It is a interesting problem.
That's a key piece of the article. He 'trusts' Gemini to classify the posts, and never hand validates anything.
And he continues building on top of that shakey trust. Is this a good idea?
It just depends on what you're doing: sounds like he's just having fun with a hobby, so it's harmless.
All he's really done is work more efficiently to save token cost and time. But he hasn't validated anything, so there's no telling if he's wasting time or not. It's a hobby though....
> So I scrape the Reddit threads where people argue about them and pull out every brand, model and steel they mention, to see what is getting bought and argued about.
Funny, normal harness with web search would one shot this, after like 20 minute search. "Replacing gemini" usually means instaling some(any)thing else, not digging deeper to get out of hole called gemini.
As for knifes, it is all same. Just do not buy total junk. Japanese knifes are way way overpriced.
This article was written with the assistance of AI. If that bothers you, stop reading here. The numbers are real: every score comes from the ten training runs described below, and the full run log is in the linked knife.day write-up.
"I didn't write any of this, but you should still trust that the remaining work, ideas and observations are all mine."
To include a disclaimer like this is to fail to recognize that "real numbers" are way less meaningful when there's clear evidence that the prompter of the LLM is not really qualified to validate them.
And this can be said to be misleading in my opinion
If he wanted a Gemini replacement verbatim, its called locally inferring it's sibling, Gemma.
And he continues building on top of that shakey trust. Is this a good idea?
It just depends on what you're doing: sounds like he's just having fun with a hobby, so it's harmless.
All he's really done is work more efficiently to save token cost and time. But he hasn't validated anything, so there's no telling if he's wasting time or not. It's a hobby though....
Funny, normal harness with web search would one shot this, after like 20 minute search. "Replacing gemini" usually means instaling some(any)thing else, not digging deeper to get out of hole called gemini.
As for knifes, it is all same. Just do not buy total junk. Japanese knifes are way way overpriced.
Genuinely appreciate the honesty. If you believe there’s nothing wrong about writing with AI, there’s no reason to not own up to it.
Okay!
To include a disclaimer like this is to fail to recognize that "real numbers" are way less meaningful when there's clear evidence that the prompter of the LLM is not really qualified to validate them.
It's a different way to think about a problem, even if it's not applicable in all areas.