英国主权AI计划用NVIDIA Nemotron支持凯尔特语言

摘要
英国主权AI计划UK-LLM基于NVIDIA Nemotron构建AI模型,支持英语和威尔士语推理。威尔士语约有85万使用者,该模型旨在赋能凯尔特语言(包括康沃尔语、爱尔兰语、苏格兰盖尔语和威尔士语)的使用者。
背景解释
凯尔特语言是英国最古老的语言,但面临数字化生存挑战。UK-LLM作为主权AI计划,通过NVIDIA技术开发多语言模型,有助于保护语言多样性,并让这些语言的使用者平等享受AI服务。此举也展示了AI在文化保护中的潜力。
原文译文
以下为抓取到的原文内容译文,已统一为站内阅读格式。

Celtic languages — including Cornish, Irish, Scottish Gaelic and Welsh — are the U.K.’s oldest living languages. To empower their speakers, the UK-LLMsovereign AIinitiative is building an AI model based onNVIDIA Nemotronthat can reason in both English and Welsh, a language spoken byabout 850,000 peoplein Wales today.
Enabling high-quality AI reasoning in Welsh will support the delivery of public services including healthcare, education and legal resources in the language.
“I want every corner of the U.K. to be able to harness the benefits of artificial intelligence. By enabling AI to reason in Welsh, we’re making sure that public services — from healthcare to education — are accessible to everyone, in the language they live by,” said U.K. Prime Minister Keir Starmer. “This is a powerful example of how the latest AI technology, trained on the U.K.’s most advanced AI supercomputer in Bristol, can serve the public good, protect cultural heritage and unlock opportunity across the country.”
The UK-LLM project, established in 2023 as BritLLM and led by University College London, has previously released two models for U.K. languages. Its new model for Welsh, developed in collaboration with Wales’ Bangor University and NVIDIA, aligns with Welsh government efforts to boost the active use of the language, with the goal of achieving a million speakers by 2050 — an initiative known asCymraeg 2050.
U.K.-based AI cloud provider Nscale will make the new model available to developers through its application programming interface. 
“The aim is to ensure that Welsh remains a living, breathing language that continues to develop with the times,” said Gruffudd Prys, senior terminologist and head of the Language Technologies Unit at Canolfan Bedwyr, the university’s center for Welsh language services, research and technology. “AI shows enormous potential to help with second-language acquisition of Welsh as well as for enabling native speakers to improve their language skills.”
This new model could also boost the accessibility of Welsh resources by enabling public institutions and businesses operating in Wales to translate content or provide bilingual chatbot services. This can help groups including healthcare providers, educators, broadcasters, retailers and restaurant owners ensure their written content is as readily available in Welsh as they are in English.
Beyond Welsh, the UK-LLM team aims to apply the same methodology used for its new model to develop AI models for other languages spoken across the U.K. such as Cornish, Irish, Scots and Scottish Gaelic — as well as work with international collaborators to build models for languages from Africa and Southeast Asia.
“This collaboration with NVIDIA and Bangor University enabled us to create new training data and train a new model in record time, accelerating our goal to build the best-ever language model for Welsh,” said Pontus Stenetorp, professor of natural language processing and deputy director for the Centre of Artificial Intelligence at University College London. “Our aim is to take the insights gained from the Welsh model and apply them to other minority languages, in the U.K. and across the globe.”
Tapping Sovereign AI Infrastructure for Model Development
The new model for Welsh is based onNVIDIA Nemotron, a family of open-source models that features open weights, datasets and recipes. The UK-LLM development team has tapped the 49-billion-parameter Llama Nemotron Super model and 9-billion-parameter Nemotron Nano model,post-trainingthem on Welsh-language data.
Compared with languages like English or Spanish, there’s less available source data in Welsh for AI training. So to create a sufficiently large Welsh training dataset, the team usedNVIDIA NIMmicroservices forgpt-oss-120bandDeepSeek-R1to translateNVIDIA Nemotron open datasetswith over 30 million entries from English to Welsh.
They used a GPU cluster through theNVIDIA DGX Cloud Leptonplatform and are harnessing hundreds ofNVIDIA GH200 Grace Hopper SuperchipsonIsambard-AI— the U.K.’s most powerful supercomputer, backed by£225 million in government investmentand based at University of Bristol — to accelerate their translation and training workloads.
This new dataset supplements existing Welsh data from the team’s previous efforts.
Capturing Linguistic Nuances With Careful Evaluation
Bangor University, located in Gwynedd — the county with thehighest percentage of Welsh speakers— is supporting the new model’s development with linguistic and cultural expertise.
Welsh translation of: “The aim is to ensure that Welsh remains a living, breathing language that continues to develop with the times.” — Gruffudd Prys, Bangor University
Prys, from the university’s Welsh-language center, brings to the collaboration about two decades of experience with language technology for Welsh. He and his team are helping to verify the accuracy of machine-translated training data and manually translated evaluation data, as well as assess how the model handles nuances of Welsh that AI typically struggles with — such as the way consonants at the beginning of Welsh words change based on neighboring words.
The model, as well as the Welsh training and evaluation datasets, are expected to be made available for enterprise and public sector use, supporting additional research, model training and application development.
“It’s one thing to have this AI capability exist in Welsh, but it’s another to make it open and accessible for everyone,” Prys said. “That subtle distinction can be the difference between this technology being used or not being used.”
Deploy Sovereign AI Models With NVIDIA Nemotron, NIM Microservices
The framework used to develop UK-LLM’s model for Welsh can serve as a foundation for multilingual AI development around the world.
Benchmark-topping Nemotron models, data and recipes are publicly available for developers to build reasoning models tailored to virtually any language, domain and workflow. Packaged as NVIDIA NIM microservices, Nemotron models are optimized for cost-effective compute and run anywhere, from laptop to cloud.
Europe’s enterprises will be able to runopen, sovereign models on the PerplexityAI-powered search engine.
_Get started with__NVIDIA Nemotron__._
- * *
_Welsh translation:_
Ymestyn Ar Draws yr Ynysoedd: Mae DU-LLM yn Dod â Deallusrwydd Artiffisial i Ieithoedd y DU Gyda NVIDIA Nemotron
_Wedi’i hyfforddi ar yr uwch gyfrifiadur Isambard-AI, mae model newydd a ddatblygwyd gan University College London, NVIDIA a Phrifysgol Bangor yn manteisio ar dechnegau a setiau data ffynhonnell agored NVIDIA Nemotron i alluogi rhesymu Deallusrwydd Artiffisial ar gyfer y Gymraeg ac ieithoedd eraill y DU ar gyfer gwasanaethau cyhoeddus gan gynnwys gofal iechyd, addysg ac adnoddau cyfreithiol._
Ieithoedd Celtaidd — gan gynnwys Cernyweg, Gwyddeleg, Gaeleg yr Alban a Chymraeg — yw ieithoedd byw hynaf y DU. Er mwyn grymuso eu siaradwyr, mae menterDeallusrwydd Artiffisial sofrany DU-LLM yn adeiladu model Deallusrwydd Artiffisial yn seiliedig arNVIDIA Nemotrona all resymu yn Saesneg a Chymraeg hefyd, iaith a siaredir gantua 850,000 o boblyng Nghymru heddiw.
Bydd galluogi rhesymu Deallusrwydd Artiffisial o ansawdd uchel yn y Gymraeg yn cefnogi’r ddarpariaeth o wasanaethau cyhoeddus gan gynnwys gofal iechyd, addysg ac adnoddau cyfreithiol yn yr iaith.
“Rwyf am i bob cwr o’r DU allu harneisio manteision deallusrwydd artiffisial. Drwy alluogi deallusrwydd artiffisial i resymu yn y Gymraeg, rydym yn sicrhau bod gwasanaethau cyhoeddus — o ofal iechyd i addysg — yn hygyrch i bawb, yn yr iaith maen nhw’n byw ynddi,” meddai Prif Weinidog y DU, Keir Starmer. “Mae hon yn enghraifft bwerus o sut y gall y dechnoleg dddiweddaraf, wedi’i hyfforddi ar uwch gyfrifiadur deallusrwydd artiffisial mwyaf datblygedig y DU ym Mryste, wasanaethu lles y cyhoedd, amddiffyn treftadaeth ddiwylliannol a datgloi cyfleoedd ledled y wlad.”
Mae prosiect DU-LLM, a sefydlwyd yn 2023 fel BritLLM ac a arweinir gan University College London, wedi rhyddhau dau fodel ar gyfer ieithoedd y DU yn flaenorol. Mae ei fodel newydd ar gyfer y Gymraeg, a ddatblygwyd mewn cydweithrediad â Phrifysgol Bangor Cymru ac NVIDIA, yn cyd-fynd ag ymdrechion llywodraeth Cymru i hybu defnydd gweithredol o’r iaith, gyda’r nod o gyflawni miliwn o siaradwyr erbyn 2050 — menter o’r enwCymraeg 2050.
Bydd darparwr cwmwl Deallusrwydd Artiffisial yn y DU, Nscale, yn sicrhau bod y model newydd ar gael i ddatblygwyr trwy ei ryngwyneb rhaglennu rhaglenni (API).
“Y nod yw sicrhau bod y Gymraeg yn parhau i fod yn iaith fyw, sy’n anadlu ac sy’n parhau i ddatblygu gyda’r oes,” meddai Gruffudd Prys, uwch derminolegydd a phennaeth yr Uned Technolegau Iaith yng Nghanolfan Bedwyr, canolfan y brifysgol ar gyfer gwasanaethau, ymchwil a thechnoleg y Gymraeg. “Mae deallusrwydd artiffisial yn dangos potensial aruthrol i helpu gyda chaffael y Gymraeg fel ail iaith yn ogystal â galluogi siaradwyr brodorol i wella eu sgiliau iaith.”
Gallai’r model newydd hwn hefyd roi hwb i hygyrchedd adnoddau Cymraeg drwy alluogi sefydliadau cyhoeddus a busnesau sy’n gweithredu yng Nghymru i gyfieithu cynnwys neu ddarparu gwasanaethau sgwrsfot dwyieithog. Gall hyn helpu grwpiau gan gynnwys darparwyr gofal iechyd, addysgwyr, darlledwyr, manwerthwyr a pherchnogion bwytai i sicrhau bod eu cynnwys ysgrifenedig yr un mor hawdd ar gael yn y Gymraeg ag y mae yn Saesneg.
Y tu hwnt i’r Gymraeg, mae tîm y DU-LLM yn anelu at gymhwyso’r un fethodoleg a ddefnyddiwyd ar gyfer ei fodel newydd i ddatblygu modelau Deallusrwydd Artiffisial ar gyfer ieithoedd eraill a siaredir ledled y DU fel Cernyweg, Gwyddeleg, Sgoteg a Gaeleg yr Alban — yn ogystal â gweithio gyda chydweithwyr rhyngwladol i adeiladu modelau ar gyfer ieithoedd o Affrica a De-ddwyrain Asia.
“Mae’r cydweithrediad hwn gydag NVIDIA a Phrifysgol Bangor wedi ein galluogi i greu data hyfforddi newydd a hyfforddi model newydd mewn amser record, gan gyflymu ein nod o adeiladu’r model iaith gorau erioed ar gyfer y Gymraeg,” meddai Pontus Stenetorp, yr athro prosesu iaith naturiol a dirprwy gyfarwyddwr y Ganolfan Deallusrwydd Artiffisial yn University College London. “Ein nod yw cymryd y mewnwelediadau a gafwyd o’r model Cymraeg a’u cymhwyso i ieithoedd lleiafrifol eraill, yn y DU ac ar draws y byd.
Manteisio ar Seilwaith Deallusrwydd Artiffisial Sofran ar gyfer Datblygu Model
Mae’r model newydd ar gyfer y Gymraeg yn seiliedig arNVIDIA Nemotron, teulu o fodelau ffynhonnell agored sy’n cynnwys pwysau, setiau data a ryseitiau agored.Mae’r tîm datblygu DU-LLM wedi manteisio ar fodel 49-biliwn-paramedr Llama Nemotron Super a model 9-biliwn-paramedr Nemotron Nano, gan euhôl hyfforddiar ddata iaith Gymraeg.
O’i gymharu ag ieithoedd fel Saesneg neu Sbaeneg, mae llai o ddata ffynhonnell ar gael yn y Gymraeg ar gyfer hyfforddiant Deallusrwydd Artiffisial. Felly, er mwyn creu set ddata hyfforddi Cymraeg ddigon mawr, defnyddiodd y tîm ficrowasanaethauNVIDIA NIMar gyfergpt-oss-120baDeepSeek-R1i gyfieithusetiau data agored NVIDIAgyda dros 30 miliwn o gofnodion o’r Saesneg i’r Gymraeg.
Defnyddion nhw glwstwr GPU drwy blatfformNVIDIA DGX Cloud Leptonac yn harneisio cannoedd oUwchsglodion NVIDIA GH200 Grace HopperarIsambard-AI— uwchgyfrifiadur mwyaf pwerus y DU, gyda chefnogaeth£225 miliwn o fuddsoddiad gan y llywodraethac wedi’i leoli ym Mhrifysgol Bryste — i gyflymu eu llwythi gwaith cyfieithu a hyfforddi.
Mae’r set ddata newydd hon yn ategu data presennol yr iaith Gymraeg o ymdrechion blaenorol y tîm.
Cipio Naws Ieithyddol Gyda Gwerthusiad Gofalus
Mae Prifysgol Bangor, sydd wedi’i lleoli yng Ngwynedd — y sir gyda’rganran uchaf o siaradwyr Cymraegs— yn cefnogi datblygiad y model newydd gydag arbenigedd ieithyddol a diwylliannol.
Mae Prys, o ganolfan Gymraeg y brifysgol, yn dod â thua dau ddegawd o brofiad gyda thechnoleg iaith ar gyfer y Gymraeg i’r cydweithrediad. Mae ef a’i dîm yn helpu i wirio cywirdeb data hyfforddi a gyfieithir gan beiriannau a data gwerthuso a gyfieithir â llaw, yn ogystal ag asesu sut mae’r model yn ymdrin â naws Gymraeg y mae Deallusrwydd Artiffisial fel arfer yn cael trafferth â nhw — megis y ffordd y mae cytseiniaid ar ddechrau geiriau Cymraeg yn newid yn seiliedig ar eiriau cyfagos.
Disgwylir i’r model, yn ogystal â’r setiau data hyfforddiant a gwerthuso’r Gymraeg, fod ar gael i fentrau a’r sector cyhoeddus eu defnyddio, gan gefnogi ymchwil ychwanegol, hyfforddiant modelu a datblygu rhaglenni.
“Mae’n un peth cael y gallu Deallusrwydd Artiffisial hwn yn bodoli yn y Gymraeg, ond mae’n beth arall ei wneud yn agored ac yn hygyrch i bawb,” meddai Prys. “Gall y gwahaniaeth cynnil hwnnw fod y gwahaniaeth rhwng y dechnoleg hon yn cael ei defnyddio ai peidio.”
Defnyddio Modelau Deallusrwydd Artiffisial Sofran Gyda NVIDIA Nemotron, Microwasanaethau NIM
Gall y fframwaith a ddefnyddiwyd i ddatblygu model DU-LLM ar gyfer y Gymraeg fod yn sylfaen ar gyfer datblygu Deallusrwydd Artiffisial amlieithog ledled y byd.
Mae modelau, data a ryseitiau Nemotron, sy’n cyrraedd y brig, ar gael yn gyhoeddus i ddatblygwyr er mwyn iddynt adeiladu modelau rhesymu sydd wedi’u teilwra i bron unrhyw iaith, parth a llif gwaith. Wedi’u pecynnu fel microgwasanaethau NVIDIA NIM, mae modelau Nemotron wedi’u hoptimeiddio ar gyfer cyfrifiadura cost-effeithiol a rhedeg yn unrhyw le, o liniadur i’r cwmwl.
Bydd mentrau Ewrop yn gallu rhedegmodelau agored, sofran ar y peiriant chwilio Perplexitywedi’i bweru gan Ddeallusrwydd Artiffisial.
_Dewch i ddechrau arni gyda__NVIDIA Nemotron__._
- Categories:
- AI
- Deep Learning
Related News

How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost
Jun 30, 2026

NVIDIA and AWS Collaborate to Bring AI to Production at Scale
Jun 23, 2026

How Businesses Are Building Specialized AI They Can Trust
Jun 23, 2026

NVIDIA Powers Over 400 of the World’s 500 Fastest Supercomputers
Jun 23, 2026
Follow Us
'/%3e %3cpath id='NVIDIA' d='M463.9,151.934v13.127h3.707V151.934Zm-29.164-.018v13.145h3.74v-10.2l2.918.01a2.674,2.674,0,0,1,2.086.724c.586.625.826,1.632.826,3.476v5.995h3.624V157.8c0-5.183-3.3-5.882-6.536-5.882Zm35.134.018v13.127h6.013c3.2,0,4.249-.533,5.38-1.727a7.352,7.352,0,0,0,1.316-4.692,7.789,7.789,0,0,0-1.2-4.516c-1.373-1.833-3.352-2.191-6.306-2.191Zm3.677,2.858h1.594c2.312,0,3.808,1.039,3.808,3.733s-1.5,3.734-3.808,3.734h-1.594Zm-14.992-2.858-3.094,10.4-2.965-10.4h-4l4.234,13.127h5.343l4.267-13.127Zm25.749,13.127h3.708V151.935h-3.709ZM494.7,151.939l-5.177,13.117h3.656l.819-2.318h6.126l.775,2.318h3.969l-5.216-13.118Zm2.407,2.393,2.246,6.145h-4.562Z' transform='translate(-399.551 -148.155)'/%3e %3cpath id='Eye_Mark' data-name='Eye Mark' d='M129.832,124.085v-1.807c.175-.013.353-.022.533-.028,4.941-.155,8.183,4.246,8.183,4.246s-3.5,4.863-7.255,4.863a4.553,4.553,0,0,1-1.461-.234v-5.478c1.924.232,2.31,1.082,3.467,3.01l2.572-2.169a6.81,6.81,0,0,0-5.042-2.462,9.328,9.328,0,0,0-1,.059m0-5.968v2.7c.177-.014.355-.025.533-.032,6.871-.232,11.348,5.635,11.348,5.635s-5.142,6.253-10.5,6.253a7.906,7.906,0,0,1-1.383-.122v1.668a9.1,9.1,0,0,0,1.151.075c4.985,0,8.59-2.546,12.081-5.559.578.463,2.948,1.591,3.435,2.085-3.319,2.778-11.055,5.018-15.44,5.018-.423,0-.829-.026-1.228-.064v2.344h18.947v-20Zm0,13.009v1.424c-4.611-.822-5.89-5.615-5.89-5.615a9.967,9.967,0,0,1,5.89-2.85v1.563h-.007a4.424,4.424,0,0,0-3.437,1.571s.845,3.035,3.444,3.908m-8.189-4.4a11.419,11.419,0,0,1,8.189-4.449v-1.463c-6.043.485-11.277,5.6-11.277,5.6s2.964,8.569,11.277,9.354v-1.555C123.731,133.451,121.643,126.728,121.643,126.728Z' transform='translate(-118.555 -118.117)' fill='%23000000'/%3e %3c/svg%3e)United States
- Privacy Policy
- Your Privacy Choices
- Terms of Service
- Accessibility
- Corporate Policies
- Product Security
- Contact
Copyright © 2026 NVIDIA Corporation
Share on Mastodon
Enter your Mastodon instance URL (optional)Share
来源地区
United States
热度分
68
分类
算力芯片
语言
en
