Apple Research Proposes Rubric-Based Method for Training Agents Across Multiple Environments
In short: Apple researchers, with National University of Singapore collaborators, published a paper on RISED, a new training framework for LLM agents that must learn across many different interactive environments at once. Instead of relying only on scalar reward signals, RISED uses an LLM judge to tag each training rollout with rubrics describing its behaviour, then uses these rubric tags to select training data and to supply extra supervision. The authors report RISED achieves the highest average success rate across environments compared to existing approaches, ranking first or second in every tested environment.
This summary was generated automatically by AI from Apple Machine Learning's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.
What changed?
- 1New method called RISED combines rubric-based data selection with on-policy self-distillation for multi-environment agent training
- 2An LLM judge labels each rollout with predefined rubric tags shared across environments
- 3Positive rubrics guide extra token-level supervision; negative rubrics steer future rollouts away from recurring failure modes
- 4Reported to achieve the highest mean pass rate across environments and rank first or second in every individual environment tested, across multiple model backbones
Why it matters
As companies push toward agents that can handle many different tools, tasks or environments with a single model, training methods that go beyond simple pass/fail rewards become important for building more capable, generalist agents — relevant background for anyone tracking how agentic AI systems are trained and evaluated.
What it means for AI agents and contact centers
This is early-stage research on training methodology rather than a deployable product, so it has no direct, immediate application for voice agent or contact-center stacks; it is useful mainly as a signal of how the industry is improving multi-task agent training, which could eventually influence future agent frameworks or tool-calling models.
Sources
- Apple Machine LearningOfficialPrimary sourceOriginal article →„RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation“6 Oct 2026, 03:00
- Published by source
- 6 Oct 2026, 03:00
- Found by our system
- 7 Oct 2026, 05:05
- Summary generated
- 7 Oct 2026, 05:06
This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy