Showing posts with label #AIAgents. Show all posts
Showing posts with label #AIAgents. Show all posts

Wednesday, August 5, 2026

Designed for the Agentic Era: AMD Ryzen AI Halo Delivers up to 34% Faster Agent Orchestration with Local Inference

Agent adoption is exploding and in most cases agents run inference in the cloud. The implications are clear: rising Cloud API costs, sensitive data leaves the device, and the workflow only runs when connected to the network. Running inference locally flips the equation: no per-token cloud cost, data that never leaves the machine, offline and low-latency operation, and full control over the model and the agent. The question is: Can you actually run a realistic agentic workload locally.

Previously, we showed that even when inference runs in the cloud, the most important upgrade might be the CPU in your system because the agent loop (planning, routing, moving data, assembling results) is orchestration, and orchestration is CPU work that runs locally no matter where the tokens are generated. Here, we complete the story for on-device inference. When the whole agent runs locally, the chat-era instinct is to ask which system has the better GPU. But agent workflows aren’t one large matrix multiply; it’s a pipeline of orchestration and inference that can include parse, OCR, chunk, embed, retrieve, generate, validate and most of those stages do not use the GPU. Workflows shaped like that are won by a balanced system, not just the biggest accelerator.

The metrics changed and that changes everything

The chat era optimized two GPU-generation metrics: time-to-first token and tokens per second. Both describe how fast one model streams one response. Agents are evaluated differently — on completed, validated work so the KPIs shift to end-to-end completion time and cost per completed workflow. Token generation is only one stage of the workflow. Orchestrating involves preparing documents, embedding and retrieving evidence, moving data between engines, and validating output which can consume a significant amount of the end-to-end workflow completion time. High tokens per second on inference alone tells you very little about how fast the job actually finished, or what it cost. The bottom line is token generation and CPU orchestration are both stages of the pipeline but end –to-end completion time becomes the more meaningful metric for agentic workflows.

A realistic workload, measured end-to-end

To test this accurately, we created HEPA, the Hermes Executive Presentation Agent, which tasks a local agent with producing a board-ready executive presentation, a supporting memo, charts, and a cited source appendix, built from a fixed local dataset that represents a large shared team drive. HEPA measures end-to-end local agent performance on a realistic enterprise knowledge-worker workflow. The dataset is varied: relevant files mixed with redundant, old, and conflicting material, plus scanned and image-derived sources that require real OCR. The agent has to read that dataset, OCR the scanned inputs, chunk and embed everything, retrieve a bounded evidence pack, generate the deliverables, and pass deterministic quality gates. Then, it’s scored on finished, quality-gated work, not isolated model throughput.

The measured setup: 302 local sources / 801 prepared chunks; orchestrator Qwen3.6-35B-A3B (mixture-of-experts, ~3B active, 4-bit) served via llama.cpp; FastEmbed ONNX nomic-embed-text-v1.5 embeddings and real Tesseract OCR both on the CPU; identical corpus, model, and settings on both machines; both platforms completed 5/5 valid runs and 25/25 deterministic checks. In this configuration, seven of the eight pipeline stages are executed on the CPU, while token generation runs on the GPU. (Full parameters in System Configurations below.)

The result: AMD finishes the job first and at lower cost

AMD Ryzen AI Halo power by the AMD Ryzen AI Max+ 395 processor was purpose-built for agentic workflows. Compared to the NVIDIA DGX Spark Grace Blackwell (GB10), both running  Linux under the same five-run workflow, the AMD Ryzen AI Halo wins on end-to-end completion time, CPU orchestration, and cost per completed workflow.

AMD Ryzen AI Halo is 34% faster vs NVIDIA DGX Spark on CPU orchestration.1 AMD completes the CPU-side preparation in 152.1 s versus 229.9 s. The decisive component is embeddings and routing, which run on the CPU — AMD 71.5 s versus 139.6 s, a nearly 2× difference.

AMD Ryzen AI Halo completes the workflow 15% faster than NVIDA DGX Spark.2 AMD completes the validated workflow in 311.6 s versus 367.1 s for DGX Spark — 55.5 seconds faster. This is the headline agentic KPI: time to finished, quality-gated work.

AMD Ryzen AI Halo delivers a 27% lower cost per completed workflow vs NVIDIA DGX Spark.3 On a 3-year amortized hardware basis at continuous utilization (energy excluded), AMD runs about $0.0132 per workflow versus roughly $0.0182 for DGX — 27% less because it costs less to acquire and finishes more work in the same time.

Even though DGX Spark has an advantage on inference, it still loses the workflow because inference is only one stage, and the CPU-bound orchestration matters. The systems were optimally configured with the AMD system running llama.cpp on Vulkan while the NVIDIA system ran llama.cpp with CUDA and the AMD system conceded the inference stage, but it still completed the job first. 

Why AMD wins the orchestration stage

The advantage comes down to how the CPU-bound stages parallelize. Parsing, OCR (thread capped), embeddings, and routing are data-parallel work that scales with cores and threads and grows with the size of the dataset. The Ryzen AI Max+ 395 processor brings 16 “Zen 5” cores and 32 threads via SMT — 32 effective workers — plus AVX-512/VNNI vector acceleration and native x86 tooling for the document and retrieval stack. NVIDIA DGX Spark’s Grace Blackwell has 20 Arm cores with no SMT, and SVE2 rather than AVX-512. More threads and wider vectors mean more source files parsed, embedded, and routed at once — so the CPU-bound preparation that dominates an agent’s completion time simply runs faster.

One system, not one chip

Under the hood it’s one orchestrating agent driving several resident models, with the work split across the system’s engines. The CPU runs the loop and the preparation stages — parsing 300+ sources, chunking, and running the embedding and routing model that decides what evidence matters. The GPU generates the tokens. And large unified memory keeps the models and data resident, so the hand-offs between them stay efficient — the capability that makes a full on-device model portfolio practical in the first place.

The chat era asked whose GPU was better. The agent era asks a better question: whose system finishes the workflow — fastest, and at lower cost? On device, the answer is a balanced system with CPU, GPU, and unified memory optimized for completion time and cost per workflow rather than tokens per second alone. AMD Ryzen AI Halo was optimized as a balanced system rather than simply maximizing GPU performance.

A system can win the token race and still lose the job. As on-device agents grow into full model portfolios — an orchestrator, an always-on small model, embedders, rerankers, OCR, speech, and image models — the balanced-system advantage compounds. Chat was a GPU story. Agents are a system story.

#####

I found this great deal on Lazada! Check it out! 

Product Name:  AMD Ryzen 5 5500 6Cores 12Threads Desktop Processor - Box | Tray Type itw

Product Price:  ₱5,950

Discount Price:  ₱5,950

https://s.lazada.com.ph/s.ZhOmBI?c=q&t=p-i3GZDZC-sVdnu88

Wednesday, June 24, 2026

PMA Charts the Next Era of Marketing with ONE MAP at the 55th Annual Marketing Conference


MANILA, Philippines. Since 1954, the Philippine Marketing Association (PMA) is presenting what it calls its most operationally ambitious marketing conference to date. At a media launch held June 24, 2026, at the Hilton Manila, Newport World Resorts, PMA officially unveiled the 55th Annual Marketing Conference under the theme ONE MAP (Omni Network Ecosystem Marketing Agility Protocol), with the conference aiming at "Charting the Future of Marketing."

The conference is set for September 23, 2026, at Hilton Manila, Newport World Resorts. A hybrid on-demand track runs from September 24 through November 2026, extending the program to participants across the region.

ONE MAP In View

ONE MAP is a pioneering initiative that goes beyond traditional learning formats by introducing an Omni-Experience, Output-Driven Learning Journey, which is a system designed to help delegates identify where conditions are shifting and respond before the window closes. The acronym reads in two halves.

The first half, O.N.E., covers the Omni Network Ecosystem. Omni pulls every channel and touchpoint into a single integrated system. The Network layer describes the partners, platforms, and communities that surround a brand. The Ecosystem pillar ties these into a customer-engagement environment built to compound over time, where each element reinforces the next.

The second half, M.A.P., covers the Marketing Agility Protocol. Marketing refers to the brand and demand engines that drive commercial growth. Agility is the organizational capacity to pivot under volatile conditions. Protocol is the standardized playbook that makes both repeatable regardless of what disruption a team is facing.

"PMA is not merely organizing another conference. PMA is presenting ONE MAP as a movement that reimagines how professionals learn and apply knowledge. It recognizes that the future belongs not to those who simply acquire information, but to those who can transform insights into action, ideas into innovation, and learning into measurable impact,” said Mitch Ballesteros, 2026 PMA President and CEO and Founder of Exlink Management and Marketing Services. She placed ONE MAP within PMA's broader 2026 organizational theme of "ONENESS," describing the conference as the association's most defining statement yet on where Philippine marketing is headed.

A Dual-Director Structure

For the first time, the 55th Marketing Conference is led by two members from PMA’s Board of Directors. Cha Pestano-Diaz, 2026 PMA Executive Vice President and Co-Founder & Marketing Director of Artisano Studio, and Adolf Aran Jr., 2026 PMA Director for National Marketing Conference and Senior Director for Corporate Marketing Communications of National University, synergize to make this PMA flagship event a success. The dual-director structure reflects the scale of this edition and the expanded scope it carries compared to previous iterations. To support them is this year’s overall conference Chairperson Greg Banzon, Chief Operating Officer and Executive Vice President of Century Pacific Food, Inc.

The conference program was designed for participants to engage across six key pillars: Meet Up Day, Digital Experience, Business Networking, Featured Company Visits, Fellowship Gala, and the Marketing Gambit Cup. Together, these experiences are designed to provide a more comprehensive and meaningful learning journey that reflects how professionals learn, connect, and grow in today’s business landscape.


What September 23 Looks Like

The Day 1 run-of-show at Hilton Manila is built around two programmatic movements. Movement 1 covers the O.N.E. half, opening with a session on AI-agent consumer journeys, followed by a panel on creator economies, communities and retail-media partnerships, and a talk on building loyalty in purpose-driven ecosystems. 

Movement 2 covers M.A.P., addressing growth on constrained budgets for Filipino brands, a panel on pivots and budget reallocation, a talk on repeatable agility playbooks, and a live "Protocol in Action" crisis simulation where panelists walk the ONE MAP steps in real time.

The Extended Program

The true differentiation of ONE MAP lies beyond its Omni-Experience approach through turning learning inspiration into transformative execution. At the center of this commitment is the Marketing Gambit Cup, the capstone program of the ONE MAP learning journey. Designed as an industry-wide challenge, the Marketing Gambit Cup encourages participants to apply the insights, strategies, and frameworks gained throughout the conference to address real-world business challenges.

Delegates will be invited to develop and present a Marketing Agility and Business Resiliency Plan based on actual business and brand case studies provided by participating organizations and sponsors. Rather than ending their learning journey with notes and presentations, participants become active contributors to the marketing community by creating actionable solutions that can drive meaningful business outcomes. The top three winners will receive scholarship grants to the Marketing Professional Certification Review Program of the Marketing Institute of the Philippines, providing a pathway toward professional certification and continued industry leadership.

As PMA celebrates its 55th Marketing Conference, ONE MAP sets a new benchmark for professional development—one that empowers marketers not only to navigate change, but to lead it.