pipeline
predspot.pipeline
¶
Pipeline Module¶
Convenience functions to run a sensible default prediction pipeline in one call, and to generate small synthetic datasets for quick experiments.
Example
from predspot.pipeline import generate_testdata, run_prediction_pipeline crimes, study_area = generate_testdata(2000, '2019-01-01', '2020-12-31', seed=0) predictions, pipeline = run_prediction_pipeline(crimes, study_area, grid_resolution=1)
generate_testdata
¶
Generate synthetic crime events inside a rectangular study area.
A thin wrapper around generate_crimes with
three hotspots and the default temporal patterns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_points
|
int
|
Number of events. |
required |
start_time
|
str
|
First possible timestamp ( |
required |
end_time
|
str
|
Last possible timestamp ( |
required |
bounds
|
tuple
|
|
DEFAULT_BOUNDS
|
seed
|
int
|
Seed for reproducibility. |
None
|
Returns:
| Type | Description |
|---|---|
tuple
|
|
Source code in src/predspot/pipeline.py
build_default_pipeline
¶
build_default_pipeline(study_area, tfreq='M', grid_resolution=1, lags=2, bandwidth='silverman', random_state=None)
Build the default Predspot pipeline: KDE mapping, seasonal/trend/diff features, quantile scaling, RFE feature selection and a random forest.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
study_area
|
GeoDataFrame
|
Study area used to build the point grid. |
required |
tfreq
|
str
|
Time frequency ( |
'M'
|
grid_resolution
|
float
|
Grid spacing in kilometers. |
1
|
lags
|
int
|
Number of lags (and STL period) of the features. |
2
|
bandwidth
|
str or float
|
KDE bandwidth, see |
'silverman'
|
random_state
|
int
|
Seed for the estimator and shuffling. |
None
|
Returns:
| Type | Description |
|---|---|
PredictionPipeline
|
An unfitted pipeline. |
Source code in src/predspot/pipeline.py
run_prediction_pipeline
¶
run_prediction_pipeline(crime_data, study_area, crime_tags=None, time_range=None, tfreq='M', grid_resolution=1, lags=2, random_state=None)
Fit the default pipeline on crime data and forecast the next period.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
crime_data
|
DataFrame
|
Events with |
required |
study_area
|
GeoDataFrame
|
Study area boundary. |
required |
crime_tags
|
list
|
Keep only these crime types. |
None
|
time_range
|
tuple
|
|
None
|
tfreq
|
str
|
Time frequency ( |
'M'
|
grid_resolution
|
float
|
Grid spacing in kilometers. |
1
|
lags
|
int
|
Number of lags (and STL period) of the features. |
2
|
random_state
|
int
|
Seed for reproducibility. |
None
|
Returns:
| Type | Description |
|---|---|
tuple
|
|
Source code in src/predspot/pipeline.py
evaluate_pipeline
¶
Cross-validate a fitted pipeline; see PredictionPipeline.evaluate.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pipeline
|
PredictionPipeline
|
A fitted pipeline. |
required |
scoring
|
str
|
|
'r2'
|
cv
|
int
|
Number of folds. |
5
|
Returns:
| Type | Description |
|---|---|
list
|
One score per fold. |