Create and run an R solver benchmark

Create and run an R solver benchmark#

This example shows how to add an R solver in a simple benchmark using benchopt’s helpers to call R code from Python.

The benchmark objective is a simple minimization task:

\[\min_{\hat{X}} \; \mathrm{MSE}(X, \hat{X})\]

We define:

  • a Python Objective that evaluates MSE between X and X_hat;

  • a Python Dataset that generates a random matrix X;

  • two solvers:

    • Python-GD implemented in Python;

    • R-PGD implemented in R and called through rpy2.

At the end, we run the benchmark and display the comparison.

# Import example helpers to define the benchmark and
# programmatically call the CLI.
from benchopt.helpers.run_examples import ExampleBenchmark
from benchopt.helpers.run_examples import benchopt_cli
from benchopt.helpers.run_examples import EXAMPLES_ROOT

First, we define the initial Python benchmark, based on the benchmark examples/minimal_benchmark. It contains an objective.py file, a simulated dataset and a full python solver based on gradient descent.

benchmark = ExampleBenchmark(
    base="minimal_benchmark", name="r_solver",
    ignore=["custom_plot.py", "example_config.yml"]
)
benchmark
            
from benchopt import BaseObjective
import numpy as np


class Objective(BaseObjective):
    # Name of the Objective function
    name = 'Quadratic'

    # The three methods below define the links between the Dataset,
    # the Objective and the Solver.
    def set_data(self, X):
        """Set the data from a Dataset to compute the objective.

        The argument are the keys of the dictionary returned by
        ``Dataset.get_data``.
        """
        self.X = X

    def get_objective(self):
        "Returns a dict passed to ``Solver.set_objective`` method."
        return dict(X=self.X)

    def evaluate_result(self, X_hat):
        """Compute the objective value(s) given the output of a solver.

        The arguments are the keys in the dictionary returned
        by ``Solver.get_result``.
        """
        return dict(value=np.linalg.norm(self.X - X_hat))

    def get_one_result(self):
        """Return one solution for which the objective can be evaluated.

        This function is mostly used for testing and debugging purposes.
        """
        return dict(X_hat=1)
from benchopt import BaseDataset

import numpy as np


class Dataset(BaseDataset):
    # Name of the Dataset, used to select it in the CLI
    name = 'simulated'

    # ``get_data()`` is the only method a dataset should implement.
    def get_data(self):
        """Load the data for this Dataset.

        Usually, the data are either loaded from disk as arrays or Tensors,
        or a dataset/dataloader object is used to allow the models to load
        the data in more flexible forms (e.g. with mini-batches).

        The dictionary's keys are the kwargs passed to ``Objective.set_data``.
        """
        return dict(X=np.random.randn(10, 2))
from benchopt import BaseSolver
import numpy as np


class Solver(BaseSolver):
    # Name of the Solver, used to select it in the CLI
    name = 'gd'

    # By default, benchopt will evaluate the result of a method after various
    # number of iterations. Setting the sampling_strategy controls how this is
    # done. Here, we use a callback function that is called at each iteration.
    sampling_strategy = 'callback'

    # Parameters of the method, that will be tested by the benchmark.
    # Each parameter ``param_name`` will be accessible as ``self.param_name``.
    parameters = {'lr': [1e-3, 1e-2]}

    # The three methods below define the necessary methods for the Solver, to
    # get the info from the Objective, to run the method and to return a
    # result that can be evaluated by the Objective.
    def set_objective(self, X):
        """Set the info from a Objective, to run the method.

        This method is also typically used to adapt the solver's parameters to
        the data (e.g. scaling) or to initialize the algorithm.

        The kwargs are the keys of the dictionary returned by
        ``Objective.get_objective``.
        """
        self.X = X
        self.X_hat = np.zeros_like(X)

    def run(self, cb):
        """Run the actual method to benchmark.

        Here, as we use a "callback", we need to call it at each iteration to
        evaluate the result as the procedure progresses.

        The callback implements a stopping mechanism, based on the number of
        iterations, the time and the evoluation of the performances.
        """
        while cb():
            self.X_hat = self.X_hat - self.lr * (self.X_hat - self.X)

    def get_result(self):
        """Format the output of the method to be evaluated in the Objective.

        Returns a dict which is passed to ``Objective.evaluate_result`` method.
        """
        return {'X_hat': self.X_hat}
#loaded from minimal_benchmark/config.yml
plot_configs:
  Subopt. (log):
    plot_kind: objective_curve
    scale: loglog
  Runtimes:
    plot_kind: bar_chart
"


We can now add a solver in R with the same algorithm. To do this, we create a new file r_pgd.py that defines a solver calling an R function via benchopt.helpers.r_lang and rpy2.

The R code is defined in a separate file r_pgd.R, loaded from Python.

R_SOLVER = EXAMPLES_ROOT / "language_solvers" / "r_pgd.py"
R_SOLVER_PY = R_SOLVER.read_text(encoding="utf-8")
R_SOLVER_R = R_SOLVER.with_suffix(".R").read_text(encoding="utf-8")

benchmark.update(
    solvers={"r_pgd.py": R_SOLVER_PY, "r_pgd.R": R_SOLVER_R},
)
            

We now update the following files:


from pathlib import Path

from benchopt import BaseSolver

import numpy as np

# Import helpers from rpy2 and benchopt.helpers.r_lang
from benchopt.helpers.r_lang import import_func_from_r_file, converter_ctx

# Import R function defined in r_pgd.R so they can be retrieved as python
# functions using `func = robjects.r['FUNC_NAME']`
R_FILE = str(Path(__file__).with_suffix('.R'))


class Solver(BaseSolver):
    name = "R-PGD"

    install_cmd = 'conda'
    requirements = ['r-base', 'rpy2']
    sampling_strategy = 'iteration'

    parameters = {'lr': [1e-3, 1e-2]}

    def set_objective(self, X):
        self.X = X
        robjects = import_func_from_r_file(R_FILE)
        self.r_gd = robjects.r['gradient_descent']

    def run(self, n_iter):
        with converter_ctx():
            coefs = self.r_gd(
                self.X, self.lr, n_iter=n_iter
            )
            self.X_hat = np.asarray(coefs)

    def get_result(self):
        return {'X_hat': self.X_hat}
##' Functions used in GD algorithm
##'
##' @title Functions used in GD algorithm
##' @author Thomas Moreau
##' @export


# Main algorithm
gradient_descent <- function(X, lr, n_iter) {
    # --------- Initialize parameter ---------
    p <- ncol(X)
    parameters <- X * 0

    # --------- Run GD for n_iter iterations ---------
    for (i in 1:n_iter) {
        # Compute the gradient
        grad <- (parameters - X)
        # # Update the parameters
        parameters <- parameters - lr * grad
    }
    return(parameters)
}
"


To run this benchmark, we need to install solver dependencies. We use benchopt install with -s to select only this solver. If R is not available in your environment, this command can install it through conda using the solver requirements.

benchopt_cli(f"install {benchmark.benchmark_dir} -s r-pgd")
$ benchopt install temp_benchmark_uepvqf6a/r_solver -s r-pgd
Installing 'r_solver' requirements
# Install
Collecting packages:
- Quadratic: already available ✓
- R-PGD: collected ✓
 done
Installing required packages for:
- R-PGD
...Retrieving notices: ...working... done
Channels:
 - conda-forge
Platform: linux-64
Collecting package metadata (repodata.json): ...working... done
Solving environment: ...working... done

Downloading and Extracting Packages: ...working... done
Preparing transaction: ...working... done
Verifying transaction: ...working... done
Executing transaction: ...working... done
#
# To activate this environment, use
#
#     $ conda activate benchopt-docs
#
# To deactivate an active environment, use
#
#     $ conda deactivate

 done
- Checking installed packages... done





Then, we can run the benchmark and show the comparison.

benchopt_cli(f"run {benchmark.benchmark_dir} -n 20 -r 4")
$ benchopt run temp_benchmark_uepvqf6a/r_solver -n 20 -r 4
Benchopt is running!
Loading objective, datasets and solvers... done.
simulated
  |--Quadratic
No seed was specified. Selected global seed: 0
    |--gd[lr=0.001]: done (not enough run)
    |--gd[lr=0.01]: done (not enough run)
    |--R-PGD[lr=0.001]: done (not enough run)
    |--R-PGD[lr=0.01]: done (not enough run)
Saving result in: 
temp_benchmark_uepvqf6a/r_solver/outputs/benchopt_run_2026-09-21_11h13m28.parq
uet
Rendering benchmark results...
   Processing
temp_benchmark_uepvqf6a/r_solver/outputs/benchopt_run_2026-09-21_11h13m28.parq
uet
done
Writing results to
temp_benchmark_uepvqf6a/r_solver/outputs/r_solver_benchopt_run_2026-09-21_11h1
3m28.html
Writing r_solver index to
temp_benchmark_uepvqf6a/r_solver/outputs/r_solver.html





Here, you should see that the R solver and Python solver obtain similar convergence profiles, with runtime differences depending on your setup.

Total running time of the script: (1 minutes 21.919 seconds)

Gallery generated by Sphinx-Gallery