AMD Advancing AI 2026 Keynote
文章语言:
简
繁
EN
Share
Minutes
原文
会议摘要
AMD highlights advancements in AI technology, including open software stacks, hardware innovations like the Instinct Mi 350 p GPU and Creo AI system, and strategic partnerships with tech leaders. The roadmap emphasizes performance, efficiency, and security, aiming to integrate AI across various computing environments, from data centers to personal devices. AMD commits to innovation and collaboration to lead the AI revolution, focusing on compute scaling, AI capabilities, and open-source contributions.
会议速览
AMD showcases its leadership in AI by developing both high-performance GPUs and CPUs, enabling agentic AI to transition from providing answers to executing actions. This breakthrough empowers employees, enhances autonomous vehicles, and revolutionizes disease diagnosis, demonstrating the company's commitment to advancing AI technology.
AMD highlights the exponential growth in AI demand, emphasizing the shift from training to inference workloads. With a focus on agentic AI, AMD anticipates that by 2026, more compute capacity will be dedicated to running AI models than training them, reflecting the increasing integration of AI into daily life and various industries.
AI agents, with their ability to solve complex problems through multiple steps, are significantly increasing compute demand, particularly for GPUs and CPUs. This has led to a rapid acceleration in the AI accelerator market, which is now forecasted to reach $1.4 trillion by 2030, approaching the size of the current semiconductor market. The shift towards genetic AI and continuous workload evolution favors GPU programmability, making them the dominant force in the AI hardware ecosystem.
AMD forecasts a transformative impact of AI on server CPUs, projecting a 50%+ growth to over $200 billion market, emphasizing the need for AI integration across devices and the edge, aiming for a $2 trillion market in high-performance computing by 2030, stressing ecosystem collaboration.
AMD emphasizes its strategy focusing on compute leadership, open platforms, and AI applications across various sectors. It highlights its extensive partnerships, advancements in hardware like the Helios AI rack, and software investments. The company underscores its achievements in the server market, GPU adoption, and the complexity of AI compute, aiming for seamless system integration and ease of use.
Discusses the Helios system, highlighting its superior AI acceleration capabilities, including advanced CPU and GPU integration, liquid cooling, and high-speed networking. Emphasizes the system's competitive edge with 15% more compute, 50% more memory capacity and bandwidth, and its role in large-scale AI operations. Announces strong customer demand and partnerships, notably with Anthropic, deploying up to 2 GW of Helios for advanced AI workloads.
The dialogue highlights the synergy between hardware and AI advancements, emphasizing the effortless deployment of models on new platforms, the critical role of AI in enhancing engineering workflows, and the commitment to improving performance, power efficiency, and security across generations of technology. A key focus is the 34x increase in throughput for high-concurrency applications, showcasing the potential of integrated AI and hardware solutions for future innovations.
A significant focus on enhancing AI efficiency is highlighted, with advancements in reducing cost per token and increasing performance through innovations like Helios. This technology promises up to 18 times more tokens per dollar and 30% more tokens per dollar than competitors, driving strong demand from major AI labs and cloud providers. The partnership with a leading AI infrastructure entity showcases the journey towards optimizing compute power for AI applications, reflecting the industry's rapid growth and the critical role of efficient computing solutions.
AI models are evolving from chatbots to sophisticated agents, driving the need for increased compute power. The partnership with AMD on deploying advanced infrastructure, such as the Helios racks, is pivotal in optimizing AI model performance and scaling capabilities, aiming for massive deployment by the end of 2027.
A collaborative effort focuses on leveraging AI to automate system optimization, shortening the deployment path for new models on AMD GPUs. The partnership aims to enhance the entire AI ecosystem by making innovations accessible, emphasizing the need for increased computational power and efficiency in future AI infrastructure.
The dialogue highlights the importance of co-designing data center systems for AI, emphasizing the need for a holistic approach that includes CPUs, GPUs, memory, networking, storage, power distribution, and cooling systems. The speaker appreciates the collaborative efforts with their team, which have led to successful productization and a focus on using AI to enhance future designs.
AMD's latest 6th Gen EPYC Venice CPU family is designed to address the evolving demands of AI, cloud, and enterprise workloads. With a focus on performance, efficiency, and total cost of ownership, the Venice family includes variants optimized for AI host nodes, agent sandboxes, and general-purpose servers. Built on TSMC's 2nm process, Venice offers up to 512 threads per socket, doubling memory and IO bandwidth, and introduces innovations like 3D V-Cache stacking for high-performance computing. This comprehensive CPU portfolio underscores AMD's leadership in data center technology.
Epic's Venice CPU outperforms competitors in AI workloads, delivering higher token processing speed, more agents per watt, and leading in core density and performance. It excels in both GPU and CPU servers, offering unmatched efficiency and software compatibility, with major OEMs and cloud providers set for Q4 rollouts.
The dialogue highlights the exponential growth in demand for AI-driven services, emphasizing the need for integrated data center systems. It discusses Meta's Compute Initiative, focusing on end-to-end system design involving hardware, networking, cooling, and power. The conversation underscores the importance of collaboration with partners like AMD for co-designing and co-creating flexible, open, and heterogeneous systems. A long-standing partnership is noted, spanning multiple CPU generations, reflecting the dynamic evolution of technology.
Discusses the growing importance of CPU and GPU synergy in handling complex workloads, highlighting advancements from Mi 300 to Mi 450, emphasizing collaborative engineering, co-design, and deployment strategies for enhanced performance and efficiency in computing systems.
Speakers emphasize the importance of co-designing AI systems early to address future challenges, focusing on integrated CPU-GPU systems, long lead times for data centers and silicon, and the need for collaborative design to unlock system potential.
A discussion on the evolving inference market, highlighting the need for tailored compute solutions for different workloads, with a focus on achieving ultra low latency through disaggregated inference, featuring collaboration between two companies to address market demands.
Cerebros, a leader in high-performance chip technology, partners with AMD to deliver an unparalleled ultra low latency inference solution. Combining AMD CPUs with Cerebros' wafer scale engine, the new system offers unmatched speed and throughput, set to revolutionize the AI market and provide customers with flexible, high-performance options for on-premise and cloud deployments.
AMD's senior vice president discusses advancements in software, highlighting the launch of Rocom AI, an agentic AI platform designed to simplify GPU programming, boost developer efficiency, and integrate AI-assisted coding into the AMD ecosystem, reflecting a significant shift in GPU software development and optimization.
The Open Software Foundation introduces an AI-assisted layer, Hyperloop, enhancing developer efficiency. Hyperloop optimizes workloads, tunes configurations, and iterates for performance, exemplified by optimizing 14 models. An AI-native interface supports developers, streamlining racom integration and enabling coding agents to understand the core stack natively.
The dialogue showcases how AI agents, like Rockham AI, simplify developer tasks by optimizing models and kernels, resulting in significant performance gains. With a focus on real-world results, the Mi 455 platform demonstrates unmatched compute and memory capabilities, supported by ecosystem partners, delivering a streamlined and powerful development experience.
The dialogue highlights the seamless deployment of a large-scale AI model using advanced AI tools, emphasizing the shift towards automated model deployment and the pivotal role of collaboration in advancing AI software development. It showcases the efficiency and innovation brought by integrating AI into the development process, exemplified by a demonstration of model deployment and a creative output, alongside the importance of partnerships in driving AI progress.
The dialogue discusses the evolution of GPU software development, highlighting the use of Triton for speed and Gluon for hyper-optimization, emphasizing the need for tailored solutions to meet varying application requirements.
A partnership highlights advancements in AI-assisted kernel generation for AMD GPUs, emphasizing the benefits of open-source collaboration in enhancing performance and efficiency across hardware and software ecosystems.
The dialogue highlights the integration of AI-assisted development with high-performance computing (HPC) and scientific computing through a modular chiplet architecture, emphasizing the Mi 430x's leadership in FP 64 performance for both AI and HPC, and the strategic importance of open-source software and abstracted interfaces in enabling AI to program complex hardware systems.
AMD highlights its strategy to deliver optimal compute solutions across various AI deployment models, from data centers to personal devices, emphasizing its CPU franchise's role in driving AI advancements and addressing unique enterprise requirements.
AMD highlights its leadership in enterprise computing with tailored solutions spanning diverse workloads and AI integration. The company emphasizes the importance of flexible CPU options to meet specific workload demands and addresses AI deployment challenges, focusing on cost predictability, data security, and seamless integration with existing infrastructure. AMD's approach supports distributed AI models, considering constraints like power, cooling, space, and ease of deployment, ensuring solutions are optimized for enterprise needs.
The Instinct Mi 350 p GPU is designed to integrate AI capabilities into existing enterprise data centers without requiring infrastructure upgrades, offering high performance and cost efficiency. It supports up to 260 billion parameters, delivering four times more tokens per second per dollar than competitors, and significantly enhances productivity across various enterprise AI workloads.
Discusses AMD's implementation of AI technologies within its own data centers, focusing on autonomous threat detection and personalized AI assistants. Highlights the benefits of intelligent routing, achieving 43% token cost reduction and 3x faster response times. Emphasizes the importance of a comprehensive ecosystem for enterprise AI deployment, including software frameworks, ISV support, and trusted partnerships. Aims to assist customers in deploying AI where it generates maximum value, whether in the cloud, on-premise, or at the edge.
A discussion on At&T's status as a leading global communications company, highlighted by a personal connection to the East Bay, emphasizing gratitude for the opportunity to reconnect with the community.
Discusses scaling AI at enterprise level, emphasizing open-source models, data sovereignty, and AMD partnership. Highlights over 100 AI models in production, reducing token costs, and launching advanced open-source AI models. Stress on integrating AI across workflows and training models for broader industry impact.
The dialogue emphasizes extending AI leadership from data centers to personal and physical intelligence, highlighting the transformative potential of personal agents in enhancing productivity and the need for robust compute infrastructure to support this evolution.
The dialogue discusses the evolution of AI models, emphasizing the shift towards smaller, more capable models that can run locally on personal devices, reducing infrastructure costs and enhancing data privacy. It highlights advancements in AI hardware, such as the AI Halo, which supports models up to 200 billion parameters, and partnerships with platforms like Hugging Face to optimize open models for real applications, streamlining the path from development to deployment.
AMD highlights advancements in AI technology, showcasing increased unified memory and parameter support in their latest Halo box. They emphasize the broad availability of personal AI systems across various platforms, facilitated by open software stacks and partnerships. The focus shifts to scaling AI from individual developers to enterprise-wide deployment, with Cisco's endorsement of this vision.
Discusses the shift in inferencing patterns, the emergence of deskside computing, and the importance of network bandwidth and token cost containment in deploying intelligence across enterprises. Highlights the collaboration between two entities to ensure governance, management, and security in the new era of distributed AI.
The dialogue explores the development of a comprehensive, secure, and scalable framework for deploying AI agents in enterprises. Key components include an isolated agent sandbox, intelligent routing, core integrations, security policy enforcement, and a unified management plane. The framework ensures end-to-end resilience, observability of agent behavior, and adherence to tokenomics guidelines, facilitating confident large-scale deployment.
Cisco Cloud Control offers comprehensive management for AI inference, enabling visibility and control over devices across desktops, data centers, and cloud environments. It ensures secure monitoring, cost containment, and scalability, with plans for general availability in the US by early fall.
AMD introduces the Creo AI system on module, designed for real-time perception, reasoning, and motion control in robots, aiming to enhance human capabilities across various industries, from surgery to agriculture, by ensuring safety, reliability, and efficiency in physical AI applications.
AMD reveals its strategy to lead the AI revolution with a robust product portfolio, including next-generation CPUs like Florence and Ravenna, and GPUs such as Mi 500 and Mi 600, designed for enhanced performance and scalability. The company emphasizes its commitment to partnerships and innovation, aiming to make AI impactful across all industries.
要点回答
Q:What is the next wave of AI expected to achieve?
A:The next wave of AI is expected to evolve beyond just thinking to taking action, with AI moving from answers to action by pushing inference compute to a tremendous scale.
Q:What is AMD's mission in the context of AI computing?
A:AMD's mission is to push the boundaries of high-performance AI computing to help solve the world's most important challenges.
Q:How is AI currently impacting various industries?
A:AI is helping researchers identify new drug candidates faster, solving previously unsolvable scientific problems, and changing the way work is done across every industry, with more capabilities being developed as AI progresses.
Q:How has the demand for AI compute power grown in a short period?
A:In just two years, the demand for AI compute power has grown nearly 160 times, with more than 35 quadrillion tokens consumed each month.
Q:What shift is expected in the use of AI compute capacity?
A:It is expected that for the first time, the world will use more AI compute to run models than to train them, with roughly 60% of global AI compute capacity being used for inference this year.
Q:What is 'agent-centric AI' and how is it affecting the compute industry?
A:'Agent-centric AI' refers to the next big step for AI where agents can answer full questions and solve problems until completion, leading to a significant increase in compute demand as these agents require extensive reasoning and data access.
Q:How is the AI accelerator market forecast to grow?
A:The AI accelerator market is expected to grow significantly, with a forecast from 500 billion by 2028 to 1.4 trillion by 2030, closely approaching the size of the entire semiconductor market.
Q:What is the projected growth for the server CPU market due to AI?
A:The server CPU market is expected to grow by over 50%, starting from a $25 billion market to over $200 billion, driven by the rapid adoption of AI and the need for more CPU infrastructure.
Q:What is the role of AI at the edge in AMD's strategy?
A:AI at the edge is a key part of AMD's strategy, with a focus on bringing intelligence to devices that sense and act in real time, alongside the cloud, to fulfill the need for AI to be infused everywhere.
Q:How is AMD's product portfolio and partnerships positioned for future success?
A:AMD is positioned for future success with the broadest product portfolio, the strongest roadmaps, and deep partnerships with companies building the future. This allows for a multi-year strategy aligned with compute leadership, open platforms, and AI integration across products.
Q:What are the specifications of the new AI accelerator developed by AMD?
A:The new AI accelerator built with TSMC's leading 2nm and 3nm process technology features 320 billion transistors and combines 12 compute and IO chiplets with 432 GB of memory, all connected by industry-leading 3D chiplets technology. It is the highest performance AI accelerator in the industry.
Q:What does the CPU board in the Helios compute tray include?
A:The CPU board in the Helios compute tray houses a high-speed 96th EPYC processor, DDR 5 memory, and all the necessary IO to feed the GPUs, along with the Sela DPU, a critical part of the networking infrastructure for front-end and scale-out bandwidth delivery.
Q:What networking capabilities does the volcano AI node provide?
A:The volcano AI node allows for up to 6 volcano nicks on a board, using the open alternative Ethernet standard. Each Helios compute tray includes two volcano boards for scale-out and scale-up networking. Additionally, there are 6 dedicated networking bricks handling scale-up networking, connecting 75 GPUs together with UA link over Ethernet with silicon from partners, demonstrating an open ecosystem approach.
Q:How does Helios compare to competitors in terms of performance?
A:Helios delivers 15% more compute, 50% more Hpm form memory capacity and memory bandwidth, and 50% more scale out bandwidth compared to the competition. This means Helios can deliver more performance for large models, more capacity for longer context, and the bandwidth to scale across thousands of GPUs.
Q:What is the physical size and technology capacity of the production hardware being deployed?
A:The production hardware weighs more than 160 pounds and stands less than 2 inches tall. It comprises more than 1,000 CDNA 5 GPU compute units, over 4,600 Ze 6 CPU cores, and 31 TB of HBM4 memory, all within a single rack. This amount of technology is required to run genetic AI at scale.
Q:When are shipments of Helios expected to start?
A:Helios is expected to start shipping at the end of the third quarter and to ramp up into the fourth quarter and the second half of the year. Customer demand for Helios is extremely strong, and it is being adopted by leading AI labs.
Q:What is Anthropic's compute strategy and how is it evolving?
A:Anthropic's compute strategy is to ensure they have the best chips for the best workloads as the industry scale grows. They have been working to optimize their use of different chips and are excited about the performance and ease of use of Helios, which they plan to deploy up to 2 GW of.
Q:Why is the deployment of up to 2 GW of Helios significant?
A:The deployment of up to 2 GW of Helios is significant because Helios is an amazing machine that is performing exceptionally well, as evidenced by an engineer being able to set it up and see performance improvements in their leading model over the weekend. The deployment also reflects a strong belief in the capabilities of AMD's architecture to service the market.
Q:What is the relationship between AMD and Anthropic regarding AI development?
A:AMD and Anthropic have a strong collaboration where they are working together to familiarize leading foundational models with AMD architecture, with a belief that this will improve services to the overall market. They are particularly excited about working on highly engineering-specific workloads to further advance AI technology.
Q:What are the key areas of excitement and potential collaboration between AMD and Anthropic?
A:Key areas of excitement and potential collaboration between AMD and Anthropic include working together for scale-up to deliver better compute results and models that add value to downstream users. Another area is security, where both companies can work to ensure safety at the chip layer, server level, and throughout the network, benefiting not just themselves but the entire industry.
Q:What performance and cost benefits does Helios provide according to system results?
A:Helios provides significant performance benefits, with up to 34 times more throughput at high concurrency compared to the prior generation. It also reduces the cost per token by up to 18 times compared to previous generations. In terms of system performance and cost, Helios offers up to 10% to 15% more performance than the competition while also providing up to 30% more tokens per dollar, resulting in a strong overall value proposition.
Q:What has been the progression of AMD's collaboration with the speaker's company?
A:The collaboration with AMD started on Mi 300 and expanded to 355. The team optimized the software stack, networking, and deployed models on AMD infrastructure. Recently, they started working with Helios racks, with engineers from both teams working side by side to optimize the software stack and run GPT-class workloads. They anticipate deploying Helios at a massive scale starting towards the end of the year and accelerating throughout 2027.
Q:How is AI expected to change the future of computing and how is AMD involved in this?
A:AI is expected to recursively improve the systems it runs on and eventually do research itself, reducing the time expert engineers spend on optimizing and enabling more efficient model deployment. AMD is involved by working with researchers to develop AI solutions that can automate significant portions of the process of programming to an AMD GPU, thus speeding up the transition from new model development to production deployment.
Q:What is the significance of the new class of infrastructure being created by AI?
A:The rise of AI is creating a new class of infrastructure that places greater importance on computing power, as reflected in the development of Venice, which is designed for this emergent need. It involves different types of server computing with varying workloads - GPU servers, AI genic servers, and traditional general-purpose servers. Having the right CPU for each workload is essential, and Venice is positioned to lead across all three classes of workloads.
Q:What are the performance and efficiency improvements of the 6th generation of Epic CPUs?
A:The 6th generation of Epic CPUs, including the Z 6 core with higher IPC and frequency, delivers up to 1.8 times more performance than the previous generation. These CPUs are built on TSMC's newest 2nm process and can support up to 512 threads per socket, offering the highest compute density while doubling memory and IO bandwidth. The chiplet architecture of Venice allows for an entire family of chips that cater to different needs, from high-frequency cores for AI host nodes to more general-purpose designs. This versatility and performance make the 6th generation of Epic CPUs the best in the data center.
Q:What is the advantage of having a consistent CPU architecture for different types of workloads?
A:Having a consistent CPU architecture enables the development of a range of chips that can cater to different types of workloads, as seen in the Epic CPU family. This approach allows for optimization across different workloads like AI, general-purpose servers, and high-performance computing. The versatility and the ability to scale performance and efficiency make this architecture unique in the industry, and it underpins the leadership of the 6th generation Epic CPUs.
Q:What is the significance of optimizing data centers and racks for genetic AI?
A:Optimizing data centers and racks for genetic AI is significant because it allows for more agents per rack, thereby providing customers with the choice to pick the right cooling, space requirements, and overall system design that best suits their data center needs.
Q:What are the details of the production status and customer demand for Venice?
A:Venice is in full production and is experiencing incredible customer demand, which is the strongest for any new EPIC generation.
Q:Which companies are expected to roll out Venice in the fourth quarter?
A:Every major server OEM and every major cloud provider are on track to begin rolling out Venice in the fourth quarter.
Q:Why is there a need to consider data centers as integrated systems?
A:There is a need to consider data centers as integrated systems because the demand for inference training and recommendation systems is increasing exponentially, and there is a vision to deliver personal superintelligence to billions of people. Coordinating the entire system, including servers, hardware, networking, cooling, and power, is crucial as none of these components are optional.
Q:What opportunities does the shift in workload and demand present for collaboration?
A:The shift in workload and demand presents opportunities for collaboration because there is a need for flexibility and heterogeneity in the industry. No single company or partner can address all needs, so collaboration is key, especially with partners like AMD, to co-create and design systems.
Q:How does the partnership between Meta and AMD contribute to innovation in CPU infrastructure?
A:The partnership between Meta and AMD contributes to innovation in CPU infrastructure through long-term collaboration, deployment of multiple generations of CPUs and GPUs, and development of Rescale OCP systems together. The partnership spans multiple products, including CPUs and GPUs, and is characterized by ongoing co-design and co-creation efforts.
Q:What are the characteristics of the different inference workloads and how does disaggregated inference address them?
A:The characteristics of different inference workloads include varying needs for throughput, latency, and compute power. Some workloads require maximum throughput or low cost, others require balanced throughput, and a new class demands very fast results. Disaggregated inference addresses these workloads by allowing independent tuning of compute and memory bandwidth requirements and delivering both great compute and great memory bandwidth.
Q:What is the role of Cerebras in providing ultra low latency inference solutions?
A:Cerebras provides ultra low latency inference solutions through the development of the world's largest and fastest chip, which is then packaged into systems, racks, and clusters. This solution is designed to meet the needs of customers who require rapid results and operates in conjunction with AMD CPUs and Cerebras' Helios and Wafer Scale Engine to create a disaggregated solution.
Q:What are the characteristics of the new solution mentioned in the speech?
A:The new solution combines the industry leaders in performance and memory capacity (Helios rack) with the highest SRAM and memory bandwidth capabilities, which allows for unmatched performance and capacity in the industry.
Q:When and where will the new solution be available?
A:The new solution will be first available in the cloud with the Cerebros platform and will be later released at a store near the customer.
Q:What unique feature does the new solution provide to customers?
A:The new solution gives customers the choice to build the solution that fits their needs, offering a huge step up in capabilities for the ultra low latency segment.
Q:How has the collaboration between the speaker's company and Cerebrals benefited customers?
A:The collaboration between the speaker's company and Cerebrals has led to a solution that offers five times the throughput while maintaining extraordinary speed, which is very compelling to customers.
Q:What is the significance of AI in transforming software development for GPUs?
A:AI is transforming software development for GPUs by reducing the time to bring up workloads, making optimization automated, and removing barriers to broad adoption. It is anticipated that AI will significantly reduce the time it takes to program GPUs and make the process more efficient.
Q:What is the purpose of the AI-assisted GPU programming platform, Racom AI?
A:Racom AI is an AI-assisted GPU programming platform that aims to simplify GPU programming for developers by allowing them to use AI coding agents to optimize their workloads, thereby improving performance and reducing complexity.
Q:How does the software stack developed by the speaker's company compare to previous versions?
A:The current software stack developed by the speaker's company, which includes features like Hyperloop, is significantly more advanced than previous versions. It is able to analyze workloads, tune configurations, select and tune kernels, adjust parallelism strategies, and iterate towards performance goals.
Q:What performance gains have been observed with the new AI-assisted tools?
A:The new AI-assisted tools, Racom AI, have resulted in performance gains of up to 38% for specific models like MiniMax M with BLM. The tools also provide up to a 2.4 times speed improvement in training performance, with innovations such as fused flash attention, more efficient checkpointing, and advanced parallelism strategies.
Q:What is the role of the OpenAI team in advancing AI development and software innovation?
A:The OpenAI team has been instrumental in pushing the limits of AI infrastructure. They have developed solutions like Triton and Gluon, which have evolved to meet the needs of different workflows and have been optimized for platforms like AMD. Their collaboration with the speaker's company has been pivotal in advancing software development and the frontier of AI.
Q:What recent collaboration involving AI and AMD GPUs has been highlighted?
A:The highlighted collaboration involves teams working closely to use AI to program AMD GPUs.
Q:What improvements have been made in AI's capability to generate high-quality GPU kernels?
A:AI has significantly improved in generating high-quality GPU kernels over the past six months.
Q:What are the two main points discussed regarding the future of AI software and collaborations?
A:The two main points are the rapid improvement in AI agents' capabilities and the benefits of AMD's open software approach to advance AI.
Q:What performance benefits does AMD's open source compiler stack provide for AI agents?
A:AMD's open source compiler stack enables AI agents to perform very low level code generation, including instruction scheduling in assembly, which has led to very significant performance gains.
Q:How does the speaker describe the capabilities and the value proposition of AI in AMD's products?
A:The speaker describes AI as picking up on the open code produced by AMD, leading to a big value proposition, particularly in creating powerful capabilities through the combination of better abstractions with AI-assisted development and hardware-software co-design.
Q:Why is the Mi 430x particularly suited for scientific computing and AI?
A:The Mi 430x is built with AMD's modular chiplet architecture, providing native FP64 hardware, which is crucial for scientific computing, and offers leadership performance across both AI and HPC workloads.
Q:What is the significance of the inflection point the industry is experiencing according to the speaker?
A:The industry is experiencing an inflection point in the ability of AI to program complex hardware systems, which is making it dramatically easier to program AI systems.
Q:What strategy is being pursued by AMD to power AI in various computing forms?
A:AMD's strategy is to deliver the right compute tangent to the right workload, continuously drive optimality across their portfolio, and support AI in data centers, enterprise, personal, and physical AI forms.
Q:What are the diverse requirements of different computing environments according to AMD?
A:Different computing environments have unique requirements: enterprises are often constrained by power and cooling, personal AI requires limited compute and memory footprint, and physical AI must perform reliably in real-world environments. AMD addresses these with a complete portfolio spanning all development models.
Q:How does the Epic portfolio accommodate the varied needs of enterprise workloads?
A:The Epic portfolio spans from 8 to 256 cores with various power, frequency, and IO capabilities, providing flexibility for customers to choose the right solution for their workload and optimize costs.
Q:What are the concerns and requirements of enterprises as they adopt AI?
A:Enterprises are concerned about predictable infrastructure costs, data security, and integrating AI into their existing infrastructure. They require a distributed deployment model encompassing a range of AI deployment environments.
Q:How is the Instinct Mi 350p designed to meet the needs of enterprises?
A:The Instinct Mi 350p is designed to fit within the power and cooling constraints of enterprise servers without requiring upgrades, supporting up to 260 billion parameters, and delivering superior performance and economics compared to competitors.
Q:What real-world performance results does the Instinct Mi 350p deliver for enterprises?
A:The Instinct Mi 350p supports up to 260 billion parameters, offers up to four times the tokens per dollar compared to the competition, and provides higher productivity with up to 2 to 5 times faster tokens per second.
Q:What kind of use cases has AMD's IT team tested with AI, and what results have they achieved?
A:AMD's IT team has tested AI in autonomous threat detection and a personalized AI assistant, reducing token costs by 43% while delivering up to 3x faster response times for local workloads.
Q:How does AMD support its customers' adoption of AI across various environments?
A:AMD supports customers with a complete ecosystem of software and solutions, including numerous AI ISVs, zero-day support for leading open models, robust frameworks, and trusted platform software and OEM partnerships.
Q:What is the overarching goal of AMD's strategy?
A:The overarching goal of AMD's strategy is to help customers deploy AI where it creates the most value, whether in the cloud, on-premise, or at the edge.
Q:What are the primary applications of AI within the enterprise mentioned by the speaker?
A:The primary applications of AI within the enterprise include fraud detection, customer care, and optimizing the placement of cell towers or Ran that is the most effective use of it inside of a network.
Q:How does the transcription of call data contribute to business insights?
A:The transcription of the over 300,000 calls received daily contributes to an enormous workload, from which insights are driven back into the business.
Q:What is the significance of the work being done with AI in various business functions?
A:The work being done with AI spans across various business functions like HR, finance, cash forecasting, and employee onboarding, which signifies the company's broader use of AI in rebuilding entire workflows.
Q:Why is it important to have world-class teams and people who think in workflows when moving AI to enterprise scale?
A:Having world-class teams and people who think in workflows is important because the process involves a lot of moving from one stage to another within a workflow, which requires well-coordinated efforts.
Q:How does the partnership with AMD benefit the company in terms of AI?
A:The partnership with AMD benefits the company by leveraging open source models, different sets of chip sets, and building and training models. The company has been able to drive down token costs by post-training a model on AMD and achieving performance and accuracy comparable to more expensive chip sets.
Q:What is the significance of the company being the first telecom company to train a model on AMD?
A:Being the first telecom company to train a model on AMD signifies a pioneering effort in using open source models and a commitment to leveraging the capabilities of AMD in the telecom industry.
Q:What is the purpose of Oel 1.0 and Oel 2.0, and why are they important?
A:The purpose of Oel 1.0 and Oel 2.0 is to provide a set of open models for the industry. Oel 1.0 was announced over a year ago with over 18 million downloads, and Oel 2.0, which is more advanced and trained on AMD, is being made available via open source to showcase the significance of training models and how they work within an enterprise.
Q:How is the company extending its leadership in enterprise computing to AI?
A:The company is extending its leadership in enterprise computing to AI by listening to customers and delivering solutions for their workloads, and by focusing on helping customers accelerate their AI journey.
Q:What is the role of personal AI in reshaping the computing landscape?
A:Personal AI plays a role in reshaping the computing landscape by extending the reach of AI beyond data centers. It involves agents that can enhance productivity and understanding in the physical world, requiring a compute ecosystem to scale together.
Q:What is the impact of personal AI on data centers and overall computing economics?
A:Personal AI impacts data centers by shifting some workloads from the cloud to PCs, reducing infrastructure costs and allowing data centers to focus on larger and more demanding tasks. It also brings advantages such as data staying with the user, applications remaining responsive, and systems working even without a network.
Q:How does the new generation of smaller AI models affect the computing landscape?
A:The new generation of smaller AI models affects the computing landscape by dramatically reducing the compute required to deliver advanced intelligence, making it possible to run workloads once reserved for the cloud on personal computing engines, thus enhancing the efficiency and optimization of AI development.
Q:What does the speaker reveal about the partnership with Hugging Face and its benefits to developers?
A:The speaker reveals that the partnership with Hugging Face will bring the OpenAI ecosystem directly onto Radeon AI Halo, providing developers with optimized performance for open models and GPT flows. This partnership benefits developers by offering them access to the right models, optimization for the platform, and the ability to turn them into real applications quickly.
Q:What is the next phase for AI according to the speaker, and what are the challenges enterprises face in deploying AI at scale?
A:The next phase for AI involves deploying AI at scale, with a fundamental shift in patterns of inferencing moving from a human-led, spiky pattern to a more persistent demand signal for infrastructure. The challenges enterprises face include security concerns, cost management, and the need to operate a large number of AI systems with the appropriate security, governance, and control.
Q:What are the critical factors to consider when deploying AI in enterprises?
A:The critical factors to consider when deploying AI in enterprises include ensuring an adequate amount of network bandwidth to support the needs of local agents, managing token costs, monitoring agent behavior for safety and security, and having a governance and management system in place.
Q:What components make up the full stack for AI deployment?
A:The full stack for AI deployment includes an isolated secure agent sandbox, intelligent routing, an MCP (likely referring to a Core Set of Integrations), security policy enforcement, resiliency in the infrastructure stack, observability of the apparatus and infrastructure, monitoring of agent behavior and tokenomics, and a unified management plane for control.
Q:How can enterprises manage their AI deployments at scale?
A:Enterprises can manage their AI deployments at scale by utilizing a management plane like Cisco Cloud Control to oversee the entire estate of AI devices, whether they are on desktops, data centers, or in the cloud. This allows for full visibility and control over the entire deployment.
Q:What are the benefits of the new AMD products for AI and robotics?
A:The new AMD products for AI and robotics, such as the AMD Creo AI system on module and the c AI robots developer platform, provide faster reactions, greater capacity for complex work, and ensure safety and control. They bring together CPU, GPU, NPU, and unified memory for simultaneous and real-time processing, and offer a complete development path from concept to real-world application.
Q:What is the future roadmap for AMD's AI and computing products?
A:AMD's future roadmap for AI and computing products includes the introduction of new CPU lines such as Florence and Ravenna, as well as continued development of their Instinct GPUs with new generations every year, maintaining a cadence of performance improvement. AMD also plans to deliver new Helios systems every year to support larger models and better total cost of ownership.






