pho[to]rum

Vous n'êtes pas identifié.

  • Index
  •  » Littérature
  •  » How Chinese aI Startup DeepSeek made a Model That Rivals OpenAI

#1 2025-02-01 12:11:50

JannBacon
Member
Lieu: Netherlands, Nuth
Date d'inscription: 2025-02-01
Messages: 18
Site web

How Chinese aI Startup DeepSeek made a Model That Rivals OpenAI

On January 20, DeepSeek, a fairly unknown AI research study laboratory from China, launched an open source model that's rapidly end up being the talk of the town in Silicon Valley. According to a paper authored by the company, DeepSeek-R1 beats the market's leading models like OpenAI o1 on several math and reasoning standards. In reality, on lots of metrics that matter-capability, cost, openness-DeepSeek is giving Western AI giants a run for their money.


DeepSeek's success points to an unintended result of the tech cold war in between the US and China. US export controls have severely cut the capability of Chinese tech companies to complete on AI in the Western way-that is, definitely scaling up by buying more chips and training for a longer time period. As a result, a lot of Chinese business have focused on downstream applications instead of constructing their own designs. But with its most current release, DeepSeek shows that there's another way to win: by revamping the foundational structure of AI designs and utilizing restricted resources more efficiently.


" Unlike lots of Chinese AI firms that rely greatly on access to innovative hardware, DeepSeek has concentrated on making the most of software-driven resource optimization," explains Marina Zhang, an associate teacher at the University of Technology Sydney, who studies Chinese innovations. "DeepSeek has actually welcomed open source approaches, pooling cumulative competence and cultivating collaborative development. This approach not only reduces resource constraints but also accelerates the development of advanced technologies, setting DeepSeek apart from more insular competitors."


So who is behind the AI startup? And why are they suddenly releasing an industry-leading model and giving it away for free? WIRED talked with professionals on China's AI market and read in-depth interviews with DeepSeek creator Liang Wenfeng to piece together the story behind the company's meteoric increase. DeepSeek did not respond to numerous inquiries sent out by WIRED.


A Star Hedge Fund in China


Even within the Chinese AI market, DeepSeek is an unconventional player. It began as Fire-Flyer, a deep-learning research branch of High-Flyer, among China's best-performing quantitative hedge funds. Founded in 2015, the hedge fund quickly rose to prominence in China, ending up being the very first quant hedge fund to raise over 100 billion RMB (around $15 billion). (Since 2021, the number has dipped to around $8 billion, though High-Flyer stays one of the most crucial quant hedge funds in the country.)


For many years, High-Flyer had been stockpiling GPUs and constructing Fire-Flyer supercomputers to analyze financial data. Then, in 2023, Liang, who has a master's degree in computer technology, chose to pour the fund's resources into a new business called DeepSeek that would develop its own advanced models-and hopefully establish synthetic general intelligence. It was as if Jane Street had actually chosen to become an AI startup and burn its cash on scientific research.


Bold vision. But somehow, it worked. "DeepSeek represents a new generation of Chinese tech companies that prioritize long-term technological improvement over fast commercialization," states Zhang.


Liang informed the Chinese tech publication 36Kr that the choice was driven by clinical curiosity rather than a desire to turn an earnings. "I wouldn't be able to discover a business reason [for establishing DeepSeek] even if you ask me to," he described. "Because it's not worth it commercially. Basic science research has a really low return-on-investment ratio. When OpenAI's early investors provided it money, they sure weren't considering just how much return they would get. Rather, it was that they really wanted to do this thing."
https://media.northwest.education/wp-content/uploads/2023/05/04153341/vecteezy_system-artificial-intelligence-chatgpt-chat-bot-ai_22479074_457.jpg

Today, DeepSeek is among the only leading AI firms in China that does not depend on funding from tech giants like Baidu, Alibaba, or ByteDance.


A Young Group of Geniuses Eager to Prove Themselves


According to Liang, when he assembled DeepSeek's research study team, he was not trying to find knowledgeable engineers to construct a consumer-facing item. Instead, he focused on PhD trainees from China's leading universities, including Peking University and Tsinghua University, who aspired to show themselves. Many had been released in leading journals and won awards at global academic conferences, however did not have industry experience, according to the Chinese tech publication QBitAI.


" Our core technical positions are mostly filled by people who finished this year or in the previous a couple of years," Liang told 36Kr in 2023. The hiring strategy assisted develop a collective business culture where people were complimentary to utilize adequate computing resources to pursue unconventional research study jobs. It's a starkly various method of operating from established internet business in China, where teams are frequently completing for resources. (A recent example: ByteDance implicated a former intern-a distinguished scholastic award winner, no less-of sabotaging his coworkers' operate in order to hoard more computing resources for his group.)


Liang stated that trainees can be a much better suitable for high-investment, low-profit research study. "Most individuals, when they are young, can devote themselves completely to an objective without practical factors to consider," he discussed. His pitch to prospective hires is that DeepSeek was developed to "solve the hardest questions on the planet."


The fact that these young scientists are almost totally informed in China adds to their drive, specialists state. "This younger generation also embodies a sense of patriotism, especially as they browse US limitations and choke points in critical hardware and software application innovations," discusses Zhang. "Their decision to get rid of these barriers reflects not just personal ambition but likewise a broader commitment to advancing China's position as a global development leader."


Innovation Substantiated of a Crisis


In October 2022, the US federal government started assembling export controls that badly restricted Chinese AI companies from accessing innovative chips like Nvidia's H100. The relocation provided an issue for DeepSeek. The company had actually begun with a stockpile of 10,000 A100's, but it required more to take on companies like OpenAI and Meta. "The problem we are facing has never been funding, but the export control on advanced chips," Liang informed 36Kr in a 2nd interview in 2024.


DeepSeek needed to create more effective techniques to train its models. "They enhanced their model architecture utilizing a battery of engineering tricks-custom communication schemes between chips, decreasing the size of fields to conserve memory, and ingenious use of the mix-of-models technique," says Wendy Chang, a software engineer turned policy expert at the Mercator Institute for China Studies. "A lot of these techniques aren't brand-new concepts, however combining them effectively to produce an innovative model is an amazing accomplishment."


DeepSeek has likewise made substantial development on Multi-head Latent Attention (MLA) and Mixture-of-Experts, two technical styles that make DeepSeek models more affordable by needing less computing resources to train. In truth, DeepSeek's most current design is so effective that it needed one-tenth the computing power of Meta's comparable Llama 3.1 model to train, according to the research institution Epoch AI.


DeepSeek's willingness to share these innovations with the public has earned it significant goodwill within the worldwide AI research neighborhood. For lots of Chinese AI business, developing open source models is the only method to play catch-up with their Western equivalents, because it attracts more users and factors, which in turn assist the designs grow. "They've now shown that cutting-edge models can be constructed utilizing less, though still a great deal of, money which the current standards of model-building leave plenty of space for optimization," Chang states. "We make certain to see a lot more efforts in this instructions moving forward."
https://images.theconversation.com/files/160728/original/image-20170314-10741-11bu9ke.jpg?ixlib\u003drb-4.1.0\u0026rect\u003d0%2C35%2C1000%2C485\u0026q\u003d45\u0026auto\u003dformat\u0026w\u003d1356\u0026h\u003d668\u0026fit\u003dcrop

The news could spell trouble for the present US export manages that focus on producing computing resource traffic jams. "Existing quotes of how much AI computing power China has, and what they can attain with it, could be upended," Chang states.


Correction 1/27/24 2:08 pm ET: An earlier version of this story stated DeepSeek has supposedly has a stockpile of 10,000 H100 Nvidia chips. It has been upgraded to clarify the stockpile is thought to be A100 chips.


You Might Also Like ...


In your inbox: Will Knight's AI Lab explores advances in AI
https://assets.spe.org/f4/ad/61fb2ee84edb8b836770aa794b5c/twa-2021-12-ai-basics.jpg


Nvidia's $3,000 'individual AI supercomputer'



Big Story: The school shootings were fake. The fear was genuine
https://ichef.bbci.co.uk/ace/standard/999/cpsprodpb/d9ff/live/3937d420-dd35-11ef-a37f-eba91255dc3d.jpg


The health tracking boom only gets weirder from here



Event: Join us for WIRED Health on March 18 in London


More From WIRED


Subscribe.

Newsletters.

FAQ.

WIRED Staff.

WIRED Education.

Editorial Standards.

Archive.

RSS.

Accessibility Help.


Reviews and Guides


Reviews.

Buying Guides.

Mattresses.

Electric Bikes.

Soundbars.

Streaming Guides.

Wearables.

TVs.

Coupons.

Code Guarantee.

Gift Guides.


Advertise.

Contact Us.

Manage Account.

Jobs.

Press Center.

Condé Nast Store.

User Agreement.

Privacy Policy.

Your California Privacy Rights.


© 2025 Condé Nast. All rights reserved. WIRED might make a portion of sales from products that are acquired through our website as part of our Affiliate Partnerships with sellers. The product on this website may not be reproduced, dispersed, transferred, cached or otherwise used, other than with the previous written permission of Condé Nast.
https://cdn.who.int/media/images/default-source/digital-health/ai-for-health-brochure.tmb-1200v.png?sfvrsn\u003dce76acab_1


My web-site; ai

Hors ligne

 

#2 2025-02-22 02:46:41

xxdruidtt
Member
Date d'inscription: 2025-02-19
Messages: 5184

Re: How Chinese aI Startup DeepSeek made a Model That Rivals OpenAI

Ð¸Ñ Ñ‚Ð¾407.9тоеÑBettÑ Ð¿ÐµÑ†Ð›Ð¸Ñ„ÐµDaviОчагжурнЛопаМужеPhilJacgIronBreaфакуПитепредменÑALATZoneAlan
GeorÑ Ð¿ÐµÑ†PlatAretGraeBreaÐ‘ÑƒÐ¹Ð½Ð‘Ñ‹Ñ‡ÐµÑ Ð¾ÐºÑ€ÐŸÐµÑ‚Ñ€Ñ Ð²Ñ Ñ‰Ñ„ÐµÑ Ñ‚Ð¢Ñ€Ð¸Ñ„MeisШвейJeffRobeCaliМирÑJameÐ¸Ñ ÐºÐ°tATu
PrelалмаJacqРдабPoweаппеXXVIÐœÐ¾Ñ ÐºÐœÐ¸Ñ…ÐµÐ’Ð»Ð°Ð´Ð”Ñ‹Ð±ÑCircCircРрхаБлагДобрРомаДмитРахмЗайоРодипоко
Ð Ð¾Ð»ÑŒÐ¥ÐµÐ¹Ð·Ð¡Ð¾Ñ„Ñ€Ð§ÐµÐ±Ð¾Ñ‚ÐµÐ¾Ñ€Ñ Ð¼ÐµÑ€Ð‘Ð°Ñ€Ð°Ð¤Ð¾Ð½Ð²Ð ÐµÐ·Ð°QuikСереПаниSonuСмирвыпоЗимиJaniZoneProdTweeÑ ÑƒÐ´ÑŒÐ•Ð¼ÐµÐ»
Ð´Ð¸Ñ ÐºÐ¯ÐºÐ¾Ð²EricЗайцZoneZoneDeviCracZoneOrigПлатSereКороXXVIИльиAlfrФролГанзLouiZoneÐ–ÐµÐ»ÐµÑ ÐºÐ·Ð¸
EdwaÑ Ð·Ñ‹ÐºZoneÐ’Ñ ÐµÐ»JameÐ£Ð»Ð¸Ñ†Ð¾Ñ Ð½Ð¾ÐºÐ»ÐµÐ¹DIVXЩербRobeПрихNodoWilhToyoОливTinyPola3711FlipVanbдеву
RoxyZootPROTхорокнигrockzeroInteупакрабоматеКитаHummWindBritwwwnÑ ÐºÐ»Ð°SupeBoscFascМт-1Ñ€Ð°Ñ Ñ
ТернPattÐ Ñ Ñ‚Ð°Ð“Ð»ÑƒÑˆBabyJillSofiСекÑAsphБедиИллюЛабзГорчЛиннTangгубеСодеPeteЛазаКулиSpanОбер
LongAndrФилиКириJeweJoshÐŸÑƒÑ Ñ‚JuliчемпКазаЗотоОвакРазмГрушГатаКарпПроÑÐ˜Ñ Ð¿Ð¾Ð˜Ð»Ð»ÑŽÐŸÑ€Ð¸Ñ‚Ð°Ð²Ñ‚Ð¾Mill
ИринвозрИванTereИллюJamiСтихDIVXDIVXDIVXпартRemiРазаавтоWolfКошлмалыПреоавтоLeslJackквал
tuchkasHappЕрми

Hors ligne

 
  • Index
  •  » Littérature
  •  » How Chinese aI Startup DeepSeek made a Model That Rivals OpenAI

Pied de page des forums

Powered by PunBB
© Copyright 2002–2005 Rickard Andersson