This example shows how to add an R solver in a simple benchmark using
benchopt’s helpers to call R code from Python.
The benchmark objective is a simple minimization task:
\[\min_{\hat{X}} \; \mathrm{MSE}(X, \hat{X})\]
We define:
a Python Objective that evaluates MSE between X and X_hat;
a Python Dataset that generates a random matrix X;
two solvers:
Python-GD implemented in Python;
R-PGD implemented in R and called through rpy2.
At the end, we run the benchmark and display the comparison.
# Import example helpers to define the benchmark and# programmatically call the CLI.frombenchopt.helpers.run_examplesimportExampleBenchmarkfrombenchopt.helpers.run_examplesimportbenchopt_clifrombenchopt.helpers.run_examplesimportEXAMPLES_ROOT
First, we define the initial Python benchmark, based on the benchmark
examples/minimal_benchmark. It contains an objective.py file,
a simulated dataset and a full python solver based on gradient descent.
frombenchoptimportBaseObjectiveimportnumpyasnpclassObjective(BaseObjective):# Name of the Objective functionname='Quadratic'# The three methods below define the links between the Dataset,# the Objective and the Solver.defset_data(self,X):"""Set the data from a Dataset to compute the objective. The argument are the keys of the dictionary returned by ``Dataset.get_data``. """self.X=Xdefget_objective(self):"Returns a dict passed to ``Solver.set_objective`` method."returndict(X=self.X)defevaluate_result(self,X_hat):"""Compute the objective value(s) given the output of a solver. The arguments are the keys in the dictionary returned by ``Solver.get_result``. """returndict(value=np.linalg.norm(self.X-X_hat))defget_one_result(self):"""Return one solution for which the objective can be evaluated. This function is mostly used for testing and debugging purposes. """returndict(X_hat=1)
frombenchoptimportBaseDatasetimportnumpyasnpclassDataset(BaseDataset):# Name of the Dataset, used to select it in the CLIname='simulated'# ``get_data()`` is the only method a dataset should implement.defget_data(self):"""Load the data for this Dataset. Usually, the data are either loaded from disk as arrays or Tensors, or a dataset/dataloader object is used to allow the models to load the data in more flexible forms (e.g. with mini-batches). The dictionary's keys are the kwargs passed to ``Objective.set_data``. """returndict(X=np.random.randn(10,2))
frombenchoptimportBaseSolverimportnumpyasnpclassSolver(BaseSolver):# Name of the Solver, used to select it in the CLIname='gd'# By default, benchopt will evaluate the result of a method after various# number of iterations. Setting the sampling_strategy controls how this is# done. Here, we use a callback function that is called at each iteration.sampling_strategy='callback'# Parameters of the method, that will be tested by the benchmark.# Each parameter ``param_name`` will be accessible as ``self.param_name``.parameters={'lr':[1e-3,1e-2]}# The three methods below define the necessary methods for the Solver, to# get the info from the Objective, to run the method and to return a# result that can be evaluated by the Objective.defset_objective(self,X):"""Set the info from a Objective, to run the method. This method is also typically used to adapt the solver's parameters to the data (e.g. scaling) or to initialize the algorithm. The kwargs are the keys of the dictionary returned by ``Objective.get_objective``. """self.X=Xself.X_hat=np.zeros_like(X)defrun(self,cb):"""Run the actual method to benchmark. Here, as we use a "callback", we need to call it at each iteration to evaluate the result as the procedure progresses. The callback implements a stopping mechanism, based on the number of iterations, the time and the evoluation of the performances. """whilecb():self.X_hat=self.X_hat-self.lr*(self.X_hat-self.X)defget_result(self):"""Format the output of the method to be evaluated in the Objective. Returns a dict which is passed to ``Objective.evaluate_result`` method. """return{'X_hat':self.X_hat}
#loaded from minimal_benchmark/config.ymlplot_configs:Subopt. (log):plot_kind:objective_curvescale:loglogRuntimes:plot_kind:bar_chart
"
We can now add a solver in R with the same algorithm.
To do this, we create a new file r_pgd.py that defines a solver calling
an R function via benchopt.helpers.r_lang and rpy2.
The R code is defined in a separate file r_pgd.R, loaded from Python.
frompathlibimportPathfrombenchoptimportBaseSolverimportnumpyasnp# Import helpers from rpy2 and benchopt.helpers.r_langfrombenchopt.helpers.r_langimportimport_func_from_r_file,converter_ctx# Import R function defined in r_pgd.R so they can be retrieved as python# functions using `func = robjects.r['FUNC_NAME']`R_FILE=str(Path(__file__).with_suffix('.R'))classSolver(BaseSolver):name="R-PGD"install_cmd='conda'requirements=['r-base','rpy2']sampling_strategy='iteration'parameters={'lr':[1e-3,1e-2]}defset_objective(self,X):self.X=Xrobjects=import_func_from_r_file(R_FILE)self.r_gd=robjects.r['gradient_descent']defrun(self,n_iter):withconverter_ctx():coefs=self.r_gd(self.X,self.lr,n_iter=n_iter)self.X_hat=np.asarray(coefs)defget_result(self):return{'X_hat':self.X_hat}
##' Functions used in GD algorithm
##'
##' @title Functions used in GD algorithm
##' @author Thomas Moreau
##' @export
# Main algorithm
gradient_descent <- function(X, lr, n_iter) {
# --------- Initialize parameter ---------
p <- ncol(X)
parameters <- X * 0
# --------- Run GD for n_iter iterations ---------
for (i in 1:n_iter) {
# Compute the gradient
grad <- (parameters - X)
# # Update the parameters
parameters <- parameters - lr * grad
}
return(parameters)
}
"
To run this benchmark, we need to install solver dependencies.
We use benchoptinstall with -s to select only this solver.
If R is not available in your environment, this command can install it
through conda using the solver requirements.