☆ Yσɠƚԋσʂ ☆
- 1.44K Posts
- 1.48K Comments
here’s a detailed of explanation of why the idea that if you respect another society then you must want to live in it is imbecilic beyond belief https://dialecticaldispatches.substack.com/p/the-myth-of-the-default-society
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•Kalashnikov Group presents Kalitka-KA anti-drone system
0·10 hours agopretty much the same thing, but with a human operator
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPtoShare Funny Videos, Images, Memes, Quotes and more @lemmy.ml•Anthropic Boasts It Would Be Profitable if You Ignore How Much It Costs to Develop AI
0·16 hours agogiven the imbalance between the spending and the profits, I don’t think tax breaks are gonna cut it
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•An excellent blogpost from Shengyu Liu, kernel engineer at DeepSeek
0·19 hours agooh haha didn’t proof read auto translate
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPtoShare Funny Videos, Images, Memes, Quotes and more @lemmy.ml•The Telegraph has a cunning plan to bring down energy prices!
0·1 day agoYes, this is a real article and they are really this stupid https://www.telegraph.co.uk/politics/2026/09/14/proscribing-houthis-bring-energy-bills-down-burnham/
☆ Yσɠƚԋσʂ ☆@lemmy.mlto
Asklemmy@lemmy.ml•What are the things/facts/developments that give you optimism or are optimistic about?
0·2 days agoPretty much any news coming out of China.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoI don’t disagree with any of that. But I think we’re talking about different things here. My point is that it’s not clear that capability will continue to scale in a useful way just because you make the model bigger. If you keep getting diminishing returns while needing vastly more resources, then it’s not economically viable to run these huge models.
So, I expect that labs focusing on more efficient architectures will outcompete those that are trying to brute force the problem. Like sure, DeepSeek isn’t small in a sense that you can run it locally, but it is small compared to other models in its class, and much more energy efficient. Whatever hardware we get down the road is going to benefit more efficient models the same way meaning that they will always have a competitive advantage.
From what I see in the latest releases from Anthropic, Fable isn’t a huge leap ahead from Opus. There is an improvement, but it’s not a definitive jump in capability the way it was from Sonnet to Opus. So, they managed to make a bigger model, but got diminishing returns, and it’s evidently so expensive to run right now that they can’t even offer it as a default.
The real progress will almost certainly be happening in hybrid architectures where people start coming up with algorithms that complement LLMs and augment their capabilities. These will be like different brain regions responsible for different tasks. For example, memory formation is an obvious example here, another would be to have a built in mathematics engine. A real huge win would be to figure out how to do few shot learning on the fly as well, for which memory is a prerequisite. So, there are plenty of things we already know that can be done much better.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoThe problem is with the context and data propagation through the network. As you keep making it bigger it becomes slower and less focused. And there is research showing that smaller models do outperform large ones on some tasks https://cacm.acm.org/news/bigger-not-necessarily-better
What I expect we’ll see going forward is more hierarchical architecture where you have finely tuned models for specific tasks with a general routing model on top. This is basically already where MoE architecture is moving now. We might also see stuff like neurosymbolics get more popular where the LLM acts as a stochastic engine within a symbolic logic system. The model can handle noisy input from the real world, and transform it into structured data that a symbolic engine can operate on.
Brute forcing the problem is a naive approach and US labs took it because they effectively had unlimited resources to train their models until now.
And when more compute becomes available, solutions that are more efficient are going to further benefit from that as well. We see this with DeepSeek right now. They focused on efficiency over capability up front, and now they have a fundamentally cheaper architecture that’s rapidly catching up in capability.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoAnd the big problem for them is that they have no leverage of Chinese labs, and it’s hard to ban open models.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoThat’s definitely a plausible option, but it’s going to be very hard to ban use of open models. They could get use of official Chinese services banned, but justifying why OpenRouter and others can’t run them is going to be a lot harder. And there’s also a ton of money invested in all these AI companies running on open models now. So, the pushback will be significant.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoThe problem for them could end up being that the economics simply don’t work. If more capable models are more power hungry, then operating them might be too expensive to justify. Or it could be that there are diminishing returns, and they simply can’t make a model that’s significantly better than the current frontier.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoI’m hoping Alibaba will start selling these things at rpi prices https://wccftech.com/alibabas-tsmc-built-5nm-risc-v-chip-xuantie-c950-now-runs-qwen-3-8-27b-model-natively-unlocking-massive-vertical-integration-tailwinds/
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoAgain, there is no reason to think that you can just keep making the model bigger and keep getting improved capability that way. In fact, we already know that’s not the case because simply making them bigger stopped being the focus. The real breakthrough is going to come from better algorithms.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoThey have no leverage over Chinese labs, and China has every incentive to continue developing this tech. The only real explanation I see here is that they’re starting to get into diminishing returns territory, investors are getting edgy, and the costs of running this stuff are exploding.
☆ Yσɠƚԋσʂ ☆@lemmy.mlOPto
Technology@lemmy.ml•So what I'm reading is that they're either making a gentleman's agreement to let Chinese labs run circles around them or the easy gains are over and they're now running into a wall
0·2 days agoOh they definitely aren’t, there’s an interview with Alibaba Cloud founder where he discusses the direction in China. Basically, the goal is to find useful niches for this tech early on, then iterate and improve. They’re not chasing AGI or trying to make one model to rule them all. That said thoough, the capabilities of Chinese models in the same domains where American ones shine are very close as well. So, I do expect that Chinese models will catch up and start surpassing American ones on their own turf before long. I’m also expecting that the trend will shift towards running smaller and local models for most things because you just don’t need a giant model to do most tasks.
















https://www.youtube.com/watch?v=cLBospQs9Hk