LLMs .what do you smoke beforehand?
LLMの品質が昨年10月頃から意図的に低下させられたと感じています。Claudeなどのモデルは以前よりも明らかに劣化し、要求されたことを避けようとしたり、簡単なタスクにも長い時間をかけたりします。結果として、LLMを使わない方がましだとさえ思えます。誰かが本当に役立つものを作っているのを見かけなくなり、良いアイデアも減りました。私が何か見落としているのでしょうか?
I think the issue is that you're using the models for tasks they're not well-suited for. For creative writing or brainstorming, they can be great, but for coding, they often produce generic or outdated solutions. You need to provide more context and be very specific about what you want.
I've noticed a regression too, but it might be due to the models being optimized for safety and alignment, which makes them more conservative and less creative. They're also getting more expensive to run, so companies might be using cheaper, less capable models under the hood.
I think the problem is that people expect LLMs to be magic. They're just tools, and like any tool, you need to learn how to use them effectively. I've found that using them for research, summarizing articles, or generating ideas works well, but you have to iterate and refine the output.
Maybe you're experiencing 'AI fatigue' – the novelty has worn off, and you're noticing the limitations more. But I've seen plenty of innovative uses of LLMs, especially in niche applications like legal document analysis or medical research. It's just that the hype has died down.
I agree that some models have degraded, but I think it's a trade-off. They're more aligned and safer, but less creative. For my work, I use a combination of different models and fine-tune them for specific tasks, which gives much better results than using a general-purpose model.
The key is to use the right tool for the job. For coding, I prefer using specialized models like Codex or Copilot, which are trained on code and produce better results. For general tasks, I use Claude or GPT-4, but I always review and edit the output. It's not perfect, but it saves time.