cvxmarkowitz¶
Motivation¶
We stand on the shoulders of CVXPY.
We solve problems arising in portfolio construction following the ideas of Harry Markowitz. Markowitz gave diversification a mathematical home in the 1950s.
Our assumption is that we solve multiple problems of the same type in a row. The input for the $n$th problem may depend on the outcome of a previous problem, e.g. the $n-1$th. Hence, we need to respect their sequential nature and order.
We can however hope that the problems we construct are DPP compliant. The first time a DPP-compliant problem is solved, CVXPY compiles it and caches the mapping from parameters to problem data. As a result, subsequent rewritings of DPP problems can be substantially faster.
In practice, the problems are not constant in size. Assets are added or removed, factors are added or removed, and so on. We expect the user is providing the number of assets a priori. We can then construct a problem suitable for a number of assets equal or smaller than the one provided. Using this approach, we keep the number of assets fixed by setting the weights for the assets not used to zero. Hence we do not need to recompile the problem as a new asset has to be added.
The padding is one-directional and it has to be consistent. Data for more
assets than the problem was built for does not fit and raises a CvxDataError,
and so does a payload whose models disagree about how large the universe is --
handing the risk model two assets and the bounds four would otherwise leave the
padded tail both riskless and unbounded, and the solver would put the whole
portfolio there. update checks this across all models before it writes
anything.
Every problem has to be constructed by a Builder. Here's a builder for a classic minimum variance problem. The builder inherits from the Builder and implements the abstract property objective. The builder remains flexible. At this stage it is possible to add or remove constraints. Only once we trigger the build() method do we construct the problem and compile it.
For injecting values for data and parameters into the problem,
we use the update method. It overwrites the
parameter values in place and returns None — there is only ever one
problem, which is precisely what lets CVXPY reuse the cached compilation.
The builder picks a risk model for you: a FactorModel when you pass factors,
a SampleCovariance otherwise. To use a different one — CVar, say — pass it in
under ModelName.RISK and the builder keeps yours instead of defaulting:
from cvxmarkowitz import MinVar
from cvxmarkowitz.names import ModelName as M
from cvxmarkowitz.risk import CVar
builder = MinVar(assets=14, model={M.RISK: CVar(alpha=0.95, rows=50, assets=14)})
print(type(builder.risk).__name__)
Installation¶
The package is not published on PyPI. Install it from the git source:
Usage¶
Build a minimum-variance problem once, then re-solve it repeatedly with new data. The compiled problem is DPP-compliant, so subsequent solves reuse the cached compilation.
import numpy as np
from cvx.linalg import cholesky
from cvxmarkowitz import MinVar
from cvxmarkowitz.names import DataNames as D
# Build a long-only, budget-constrained minimum-variance problem for 4 assets.
problem = MinVar(assets=4).build()
# Inject data and parameters. Here only 2 of the 4 asset slots are used;
# the unused assets are pinned to zero weight.
problem.update(
**{
D.CHOLESKY: cholesky(np.array([[1.0, 0.5], [0.5, 2.0]])),
D.LOWER_BOUND_ASSETS: np.zeros(2),
D.UPPER_BOUND_ASSETS: np.ones(2),
D.VOLA_UNCERTAINTY: np.zeros(2),
}
)
objective = problem.solve() # defaults to the CLARABEL solver
print("objective:", round(objective, 4))
print("weights:", np.round(problem.weights, 3))
Reusing a built problem¶
The problems are parameterized (DPP-compliant), so a single built problem can be
re-solved with fresh data. Build once, then update and solve in a loop — the
cvxpy canonicalization is paid for on the first solve and reused afterwards.
problem = MinVar(assets=4).build() # build once
for correlation in (0.0, 0.5, 0.9):
problem.update(
**{
D.CHOLESKY: cholesky(np.array([[1.0, correlation], [correlation, 2.0]])),
D.LOWER_BOUND_ASSETS: np.zeros(2),
D.UPPER_BOUND_ASSETS: np.ones(2),
D.VOLA_UNCERTAINTY: np.zeros(2),
}
)
print(f"rho={correlation}: objective={problem.solve():.4f}")
Errors¶
Everything the package raises derives from CvxError, so a single
except CvxError still catches all of it. The subclasses below separate the
failure modes that want different handling:
| Error | Raised when | Retry helps? |
|---|---|---|
CvxDataError |
required data is missing, or shapes disagree | yes, with corrected input |
CvxBuildError |
the assembled problem is not DPP-compliant | no — the formulation has to change |
CvxSolverError |
the solver returned a non-optimal status | no — try another solver or relax the problem |
The CvxBuildError check is a raise rather than an assert, so it still fires
under python -O.
Development¶
This project uses uv and a
Rhiza-managed Makefile. To create the
virtual environment defined in pyproject.toml and locked in uv.lock:
marimo¶
We install marimo on the fly within the virtual environment. Executing
will install and start marimo.
experiments¶
experiments/ holds standalone research scripts — backtests and the figures
behind the talks — not part of the installed package. They pull in the dev
dependency group (yfinance, loguru, cvxsimulator, tinycta, plotly),
so run them against the full development environment:
They are deliberately outside the quality gates: make typecheck,
make docs-coverage, make deptry and make security all scope to src/, and
the coverage gate scopes to tests/. Only make fmt reaches them, since
pre-commit runs repo-wide. Treat them as scratch work — if something here earns
a stability guarantee, it belongs in src/cvxmarkowitz/.