pdmlabs.experiment.batch.auto_profile_semi_supervised_experiment#
Classes
Semi-supervised anomaly detection with automatic profile-based learning. |
- class pdmlabs.experiment.batch.auto_profile_semi_supervised_experiment.AutoProfileSemiSupervisedPdMExperiment(*args, **kwargs)#
Bases:
PdMExperimentSemi-supervised anomaly detection with automatic profile-based learning.
This experiment flavor implements an “auto-profiling” semi-supervised approach:
For each target scenario, uses an initial profile (first N timesteps) as normal behavior
Fits the anomaly detection method only on this profile
Applies the fitted method to detect anomalies in the rest of the scenario
Automatically determines the profile size via hyperparameter search
This is useful when: - You have unlabeled data with clear patterns at the start (normal operating condition) - You want to adapt to gradual drift without constant retraining - You have limited labeled anomaly examples
The “auto-profiling” optimization searches over two independent window sizes to find the normal-behaviour window that yields the best performance:
initial_profile_size– the profile collected at the start of each target source, i.e. the slice running from the beginning of the scenario to its first reset date.profile_size– the profile re-collected after every subsequent reset.
They are separate because the first slice of a source is often the only one with plenty of clean history behind it, so it can usually afford a larger profile than the shorter segments between later resets.
- pipeline#
Must have ‘failure’ or ‘reset’ events to define scenario boundaries.
- Type:
- param_space#
Must include ‘profile_size’; ‘initial_profile_size’ is optional and falls back to ‘profile_size’ when absent. Example: {‘profile_size’: [10, 20, 50], ‘initial_profile_size’: [50, 100],
‘method_alpha’: [0.1, 0.5, 1.0]}
- Type:
dict
- Raises:
IncompatibleMethodException – If method does not implement SemiSupervisedMethodInterface.
ValueError – If pipeline lacks required event definitions.
Examples
>>> from pdmlabs.method.isolation_forest import IsolationForest >>> from pdmlabs.preprocessing.no_preprocessor import NoPreprocessor >>> # ... setup pipeline ... >>> param_space = { ... 'profile_size': [10, 20, 50], ... 'method_alpha': [0.1, 1.0] ... } >>> experiment = AutoProfileSemiSupervisedPdMExperiment( ... experiment_name='auto-profile-demo', ... pipeline=pipeline, ... param_space=param_space, ... num_iteration=30, ... n_jobs=4 ... ) >>> results = experiment.execute() >>> print(f"Best profile size: {results['best_params']['profile_size']}") Best profile size: 20
- execute() dict#
Run the auto-profile semi-supervised optimization experiment.
Searches parameter space to find the best profile size and method parameters. For each combination:
For each target scenario: a. Segments by reset/failure events b. Uses the first N timesteps as the normal pattern – N is
initial_profile_size for the first segment, profile_size thereafter
Fits method on profile
Predicts on remaining data
Applies postprocessor and thresholder
Evaluates across all scenarios using PdM metrics
Returns best parameters
- Returns:
- Result dictionary with:
’best_params’: Best found parameters (includes profile_size)
’best_objective’: Best metric value achieved
’th’: Best threshold for decision boundary
- Return type:
dict
- Raises:
IncompatibleMethodException – If method is not SemiSupervisedMethodInterface.
Exception – If pipeline setup is invalid or data processing fails.
Examples
>>> experiment = AutoProfileSemiSupervisedPdMExperiment(...) >>> results = experiment.execute() >>> print(results['best_params']['profile_size']) 25