Skip to content

ml_modelling

predspot.ml_modelling

Machine Learning Modelling Module

The prediction pipeline of Predspot and thin wrappers that make scikit-learn feature selectors and regressors keep pandas indexes, so predictions stay attached to their (t, places) labels.

FeatureSelection

FeatureSelection(estimator)

Bases: TransformerMixin, BaseEstimator

Wrap a scikit-learn feature selector so that it returns DataFrames.

Parameters:

Name Type Description Default
estimator object

A selector exposing support_ after fit (e.g. RFE).

required
Source code in src/predspot/ml_modelling.py
def __init__(self, estimator):
    self.estimator = estimator

support_ property

support_

numpy.ndarray: Boolean mask of the selected columns.

Model

Model(estimator)

Bases: RegressorMixin, BaseEstimator

Wrap a scikit-learn regressor so that predictions come back as DataFrames.

Parameters:

Name Type Description Default
estimator RegressorMixin

Any scikit-learn regressor.

required
Source code in src/predspot/ml_modelling.py
def __init__(self, estimator):
    self.estimator = estimator

PredictionPipeline

PredictionPipeline(mapping, fextraction, estimator, random_state=None)

Bases: RegressorMixin, BaseEstimator

End-to-end crime hotspot prediction.

The pipeline chains three stages: a spatio-temporal mapping (e.g. KDE) that turns events into a series per place and period; a feature extraction step (e.g. PandasFeatureUnion of lag features) and a scikit-learn estimator (or Pipeline) that learns to predict the next period's value from the features.

Parameters:

Name Type Description Default
mapping SpatioTemporalMapping required
fextraction object

A transformer taking the series and returning features.

required
estimator object

A scikit-learn regressor or Pipeline whose last step returns a DataFrame with a crime_density column (see Model).

required
random_state int

Seed used to shuffle the training rows.

None
Source code in src/predspot/ml_modelling.py
def __init__(self, mapping, fextraction, estimator, random_state=None):
    self.mapping = mapping
    self.fextraction = fextraction
    self.estimator = estimator
    self.random_state = random_state
    self._offset = tfreq_offset(mapping.tfreq)
    self._stseries = None
    self._dataset = None
    self._X = None
    self._t_plus_one = None

grid property

grid

GeoDataFrame: Spatial grid used by the mapping.

dataset property

dataset

Dataset: The dataset the pipeline was fitted on.

stseries property

stseries

pandas.Series: Spatio-temporal series (observed + predicted periods).

features property

features

pandas.DataFrame: Features of every (t, places) row.

next_time property

next_time

pandas.Timestamp: The period that the next predict call forecasts.

feature_importances property

feature_importances

Importance of each selected feature.

Works when estimator is a Pipeline whose last step exposes feature_importances_ (e.g. Model around a random forest); an optional FeatureSelection step before it is taken into account.

Returns:

Type Description
DataFrame

Importance per feature, sorted descending.

fit

fit(dataset, y=None)

Fit the mapping, the features and the estimator on a dataset.

Parameters:

Name Type Description Default
dataset Dataset

Crime events and study area.

required
y None

Ignored; present for scikit-learn compatibility.

None

Returns:

Type Description
PredictionPipeline

self.

Source code in src/predspot/ml_modelling.py
def fit(self, dataset, y=None):
    """
    Fit the mapping, the features and the estimator on a dataset.

    Args:
        dataset (predspot.Dataset): Crime events and study area.
        y (None): Ignored; present for scikit-learn compatibility.

    Returns:
        PredictionPipeline: ``self``.
    """
    logger.debug("Fitting prediction pipeline")
    self._dataset = dataset
    self._stseries = self.mapping.fit_transform(dataset.crimes)
    self._X = self.fextraction.fit_transform(self._stseries)
    t0 = self._X.index.get_level_values("t").min()
    tf = self._stseries.index.get_level_values("t").max()
    X = self._X.loc[t0:tf].sample(frac=1, random_state=self.random_state)
    y = self._stseries.loc[X.index]
    self.estimator.fit(X, y)
    self._t_plus_one = self._X.index.get_level_values("t").max()
    logger.debug("Pipeline fitted on %d rows; next period is %s", len(X), self._t_plus_one)
    return self

predict

predict()

Forecast the next period for every place.

Each call appends its forecast to the series and recomputes the features, so calling it repeatedly walks forward in time.

Returns:

Type Description
DataFrame

crime_density indexed by (t, places) for the forecast period.

Source code in src/predspot/ml_modelling.py
def predict(self):
    """
    Forecast the next period for every place.

    Each call appends its forecast to the series and recomputes the
    features, so calling it repeatedly walks forward in time.

    Returns:
        pandas.DataFrame: ``crime_density`` indexed by ``(t, places)``
        for the forecast period.
    """
    self._check_fitted()
    X = self._X.loc[[self._t_plus_one], :]
    y_pred = pd.DataFrame(self.estimator.predict(X), index=X.index)
    y_pred.columns = ["crime_density"]
    logger.debug("Predicted %d places for %s", len(y_pred), self._t_plus_one)
    self._stseries = pd.concat([self._stseries, y_pred["crime_density"]]).sort_index()
    self._stseries.name = "crime_density"
    self._X = self.fextraction.transform(self._stseries)
    self._t_plus_one = self._t_plus_one + self._offset
    return y_pred

evaluate

evaluate(scoring='r2', cv=5)

Score the estimator with time series cross-validation.

Periods are split in cv consecutive folds (sklearn.model_selection.TimeSeriesSplit); the estimator is refitted on the original data afterwards.

Parameters:

Name Type Description Default
scoring str or list

'r2', 'mse' or a list of them.

'r2'
cv int

Number of folds (must be lower than the number of periods).

5

Returns:

Type Description
list or DataFrame

One score per fold; with a list of scorings, a DataFrame with one column per scoring and one row per fold.

Source code in src/predspot/ml_modelling.py
def evaluate(self, scoring="r2", cv=5):
    """
    Score the estimator with time series cross-validation.

    Periods are split in ``cv`` consecutive folds
    (``sklearn.model_selection.TimeSeriesSplit``); the estimator is
    refitted on the original data afterwards.

    Args:
        scoring (str or list): ``'r2'``, ``'mse'`` or a list of them.
        cv (int): Number of folds (must be lower than the number of periods).

    Returns:
        list or pandas.DataFrame: One score per fold; with a list of
        scorings, a DataFrame with one column per scoring and one row
        per fold.
    """
    self._check_fitted()
    scorings = [scoring] if isinstance(scoring, str) else list(scoring)
    unknown = [s for s in scorings if s not in SCORERS]
    if unknown or not scorings:
        raise ValueError('invalid scoring. Try "r2" or "mse".')
    timestamps = (
        self._X.index.get_level_values("t")
        .unique()
        .intersection(self._stseries.index.get_level_values("t").unique())
        .sort_values()
    )
    if not isinstance(cv, int) or cv >= len(timestamps):
        raise ValueError("cv must be an integer lower than the number of periods.")
    scores = {name: [] for name in scorings}
    for train_t, test_t in TimeSeriesSplit(cv).split(timestamps):
        X_train = self._X.loc[idx[timestamps[train_t], :], :].sample(
            frac=1, random_state=self.random_state
        )
        X_test = self._X.loc[idx[timestamps[test_t], :], :]
        y_train = self._stseries.loc[X_train.index]
        y_test = self._stseries.loc[X_test.index]
        self.estimator.fit(X_train, y_train)
        y_pred = self.estimator.predict(X_test)
        for name in scorings:
            scores[name].append(SCORERS[name](y_test, y_pred))
    logger.debug("%s-fold CV scores: %s", cv, scores)
    self.fit(self._dataset)  # back to normal
    if isinstance(scoring, str):
        return scores[scoring]
    return pd.DataFrame(scores, index=[f"fold {i + 1}" for i in range(cv)])