Skip to content
Wednesday, 7 October 2026
Tenesys AI News
Subscribe

Apple Research Proposes Rubric-Based Method for Training Agents Across Multiple Environments

In short: Apple researchers, with National University of Singapore collaborators, published a paper on RISED, a new training framework for LLM agents that must learn across many different interactive environments at once. Instead of relying only on scalar reward signals, RISED uses an LLM judge to tag each training rollout with rubrics describing its behaviour, then uses these rubric tags to select training data and to supply extra supervision. The authors report RISED achieves the highest average success rate across environments compared to existing approaches, ranking first or second in every tested environment.

Source: Apple Machine LearningAppleOriginal article ↗

This summary was generated automatically by AI from Apple Machine Learning's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1New method called RISED combines rubric-based data selection with on-policy self-distillation for multi-environment agent training
  • 2An LLM judge labels each rollout with predefined rubric tags shared across environments
  • 3Positive rubrics guide extra token-level supervision; negative rubrics steer future rollouts away from recurring failure modes
  • 4Reported to achieve the highest mean pass rate across environments and rank first or second in every individual environment tested, across multiple model backbones

Why it matters

As companies push toward agents that can handle many different tools, tasks or environments with a single model, training methods that go beyond simple pass/fail rewards become important for building more capable, generalist agents — relevant background for anyone tracking how agentic AI systems are trained and evaluated.

What it means for AI agents and contact centers

This is early-stage research on training methodology rather than a deployable product, so it has no direct, immediate application for voice agent or contact-center stacks; it is useful mainly as a signal of how the industry is improving multi-task agent training, which could eventually influence future agent frameworks or tool-calling models.

Sources

  • Apple Machine LearningOfficialPrimary source
    „RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation“
    6 Oct 2026, 03:00
    Original article →
Published by source
6 Oct 2026, 03:00
Found by our system
7 Oct 2026, 05:05
Summary generated
7 Oct 2026, 05:06

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

Apple Research Proposes Rubric-Based Method for Training Agents Across Multiple Environments · TENESYS AI NEWS