NVIDIA Jetson Thor speeds edge agentic inference in MLPerf v6.1
Edge AI agents face rigorous MLPerf testing across multi-turn trajectories, where NVIDIA Jetson Thor hardware cut execution runtime 6.4x against reference software configurations.
MLPerf Inference v6.1 from MLCommons added an Edge Agentic Inference workload alongside an End-to-End Retrieval-Augmented Generation benchmark. The workload evaluates autonomous software-engineering tasks rather than separate single-turn prompts. Models must execute consecutive turns, manage tools, and parse feedback while operating within tight local power and memory limits.
Miro Hodak, MLPerf Inference working group co-chair, said: “We are working hard to ensure that the MLPerf Inference benchmark continues to reflect the scenarios that the AI community values most.
“We added the end-to-end RAG test because it’s clear that query-answering has evolved beyond simply an LLM trained on a corpus; stakeholders need to understand the real-world performance of the types of multi-step, multi-component pipelines that are being built today.
“Likewise, we added the Edge Agentic Inference test because complex...
Copyright of this story solely belongs to iottechnews.com. To see the full text click HERE