ml_modelling
predspot.ml_modelling
¶
Machine Learning Modelling Module¶
The prediction pipeline of Predspot and thin wrappers that make scikit-learn
feature selectors and regressors keep pandas indexes, so predictions stay
attached to their (t, places) labels.
FeatureSelection
¶
Bases: TransformerMixin, BaseEstimator
Wrap a scikit-learn feature selector so that it returns DataFrames.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
estimator
|
object
|
A selector exposing |
required |
Source code in src/predspot/ml_modelling.py
Model
¶
Bases: RegressorMixin, BaseEstimator
Wrap a scikit-learn regressor so that predictions come back as DataFrames.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
estimator
|
RegressorMixin
|
Any scikit-learn regressor. |
required |
Source code in src/predspot/ml_modelling.py
PredictionPipeline
¶
Bases: RegressorMixin, BaseEstimator
End-to-end crime hotspot prediction.
The pipeline chains three stages: a spatio-temporal mapping (e.g.
KDE) that turns events into a series per
place and period; a feature extraction step (e.g.
PandasFeatureUnion of lag features) and a
scikit-learn estimator (or Pipeline) that learns to predict the
next period's value from the features.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mapping
|
SpatioTemporalMapping
|
required | |
fextraction
|
object
|
A transformer taking the series and returning features. |
required |
estimator
|
object
|
A scikit-learn regressor or |
required |
random_state
|
int
|
Seed used to shuffle the training rows. |
None
|
Source code in src/predspot/ml_modelling.py
feature_importances
property
¶
Importance of each selected feature.
Works when estimator is a Pipeline whose last step exposes
feature_importances_ (e.g. Model around a random
forest); an optional FeatureSelection step
before it is
taken into account.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Importance per feature, sorted descending. |
fit
¶
Fit the mapping, the features and the estimator on a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset
|
Dataset
|
Crime events and study area. |
required |
y
|
None
|
Ignored; present for scikit-learn compatibility. |
None
|
Returns:
| Type | Description |
|---|---|
PredictionPipeline
|
|
Source code in src/predspot/ml_modelling.py
predict
¶
Forecast the next period for every place.
Each call appends its forecast to the series and recomputes the features, so calling it repeatedly walks forward in time.
Returns:
| Type | Description |
|---|---|
DataFrame
|
|
Source code in src/predspot/ml_modelling.py
evaluate
¶
Score the estimator with time series cross-validation.
Periods are split in cv consecutive folds
(sklearn.model_selection.TimeSeriesSplit); the estimator is
refitted on the original data afterwards.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scoring
|
str or list
|
|
'r2'
|
cv
|
int
|
Number of folds (must be lower than the number of periods). |
5
|
Returns:
| Type | Description |
|---|---|
list or DataFrame
|
One score per fold; with a list of scorings, a DataFrame with one column per scoring and one row per fold. |