> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/harbor-framework/harbor/llms.txt
> Use this file to discover all available pages before exploring further.

# Metrics

> Metric configuration and base classes for computing evaluation metrics

## Overview

Harbor's metric system allows you to compute aggregate statistics across trial results. Metrics process lists of rewards and produce summary values like mean, max, min, or custom computations.

## MetricConfig

Configuration for a metric to be computed.

**Import:** `from harbor.models.metric.config import MetricConfig`

### Fields

<ParamField path="type" type="MetricType" default="MetricType.MEAN">
  The type of metric to compute.
</ParamField>

<ParamField path="kwargs" type="dict[str, Any]" default="{}">
  Additional keyword arguments passed to the metric constructor.
</ParamField>

### Example

```python theme={null}
from harbor.models.metric.config import MetricConfig
from harbor.models.metric.type import MetricType
from harbor.models.job.config import JobConfig

config = JobConfig(
    job_name="metrics-example",
    agents=[{"name": "claude-code"}],
    datasets=[{"name": "terminal-bench", "version": "2.0"}],
    metrics=[
        MetricConfig(type=MetricType.MEAN),
        MetricConfig(type=MetricType.MAX),
        MetricConfig(type=MetricType.MIN),
        MetricConfig(
            type=MetricType.UV_SCRIPT,
            kwargs={"script_path": "./custom_metric.py"}
        )
    ]
)
```

## MetricType

Enum defining available metric types.

**Import:** `from harbor.models.metric.type import MetricType`

### Values

<ParamField path="MEAN" type="str" value="'mean'">
  Compute the mean (average) of rewards.
</ParamField>

<ParamField path="MAX" type="str" value="'max'">
  Compute the maximum reward value.
</ParamField>

<ParamField path="MIN" type="str" value="'min'">
  Compute the minimum reward value.
</ParamField>

<ParamField path="SUM" type="str" value="'sum'">
  Compute the sum of all rewards.
</ParamField>

<ParamField path="UV_SCRIPT" type="str" value="'uv-script'">
  Run a custom Python script to compute metrics.
</ParamField>

### Example

```python theme={null}
from harbor.models.metric.type import MetricType
from harbor.models.metric.config import MetricConfig

# Built-in metrics
mean_metric = MetricConfig(type=MetricType.MEAN)
max_metric = MetricConfig(type=MetricType.MAX)
min_metric = MetricConfig(type=MetricType.MIN)
sum_metric = MetricConfig(type=MetricType.SUM)

# Custom script metric
custom_metric = MetricConfig(
    type=MetricType.UV_SCRIPT,
    kwargs={
        "script_path": "./metrics/custom.py",
        "param1": "value1"
    }
)
```

## BaseMetric

Abstract base class for implementing custom metrics.

**Import:** `from harbor.metrics.base import BaseMetric`

### Type Parameters

<ParamField path="T" type="TypeVar">
  The type of values the metric processes (typically `dict[str, float | int]` for rewards).
</ParamField>

### Abstract Methods

#### compute

```python theme={null}
@abstractmethod
def compute(self, rewards: list[T | None]) -> dict[str, float | int]
```

Compute metric values from a list of rewards.

<ParamField path="rewards" type="list[T | None]" required>
  List of reward values. None values represent failed trials or missing rewards.
</ParamField>

<ResponseField name="metrics" type="dict[str, float | int]" required>
  Dictionary of computed metric name-value pairs.
</ResponseField>

### Example Implementation

```python theme={null}
from harbor.metrics.base import BaseMetric

class PassRateMetric(BaseMetric[dict[str, float]]):
    """Compute the pass rate (percentage of rewards == 1.0)."""
    
    def compute(self, rewards: list[dict[str, float] | None]) -> dict[str, float | int]:
        # Filter out None values
        valid_rewards = [r for r in rewards if r is not None]
        
        if not valid_rewards:
            return {"pass_rate": 0.0, "n_trials": 0}
        
        # Count how many have score == 1.0
        passes = sum(1 for r in valid_rewards if r.get("score") == 1.0)
        pass_rate = passes / len(valid_rewards)
        
        return {
            "pass_rate": pass_rate,
            "n_trials": len(valid_rewards),
            "n_passed": passes
        }

# Use the metric
metric = PassRateMetric()
rewards = [
    {"score": 1.0},
    {"score": 0.5},
    {"score": 1.0},
    None,  # Failed trial
    {"score": 0.0}
]

result = metric.compute(rewards)
print(result)
# {'pass_rate': 0.5, 'n_trials': 4, 'n_passed': 2}
```

### Built-in Metrics

Harbor includes several built-in metrics:

#### Mean

```python theme={null}
from harbor.metrics.mean import Mean

metric = Mean()
rewards = [
    {"score": 1.0, "efficiency": 0.8},
    {"score": 0.5, "efficiency": 0.9},
    {"score": 0.0, "efficiency": 0.7}
]

result = metric.compute(rewards)
# {'score': 0.5, 'efficiency': 0.8}
```

#### Max

```python theme={null}
from harbor.metrics.max import Max

metric = Max()
rewards = [
    {"score": 0.8},
    {"score": 0.95},
    {"score": 0.6}
]

result = metric.compute(rewards)
# {'score': 0.95}
```

#### Min

```python theme={null}
from harbor.metrics.min import Min

metric = Min()
rewards = [
    {"score": 0.8},
    {"score": 0.95},
    {"score": 0.6}
]

result = metric.compute(rewards)
# {'score': 0.6}
```

#### Sum

```python theme={null}
from harbor.metrics.sum import Sum

metric = Sum()
rewards = [
    {"points": 10},
    {"points": 20},
    {"points": 15}
]

result = metric.compute(rewards)
# {'points': 45}
```

## VerifierResult

Contains the verification results including rewards.

**Import:** `from harbor.models.verifier.result import VerifierResult`

### Fields

<ParamField path="rewards" type="dict[str, float | int] | None" default="None">
  Dictionary mapping reward names to numeric values. None if verification failed or was disabled.
</ParamField>

### Example

```python theme={null}
from harbor.models.verifier.result import VerifierResult

# Single reward
result = VerifierResult(
    rewards={"score": 1.0}
)

# Multiple rewards
result = VerifierResult(
    rewards={
        "correctness": 1.0,
        "efficiency": 0.85,
        "style": 0.9
    }
)

# No rewards (verification failed)
result = VerifierResult(
    rewards=None
)
```

## Writing Custom Metrics

### Simple Aggregation Metric

```python theme={null}
from harbor.metrics.base import BaseMetric

class MedianMetric(BaseMetric[dict[str, float]]):
    """Compute the median of each reward."""
    
    def compute(self, rewards: list[dict[str, float] | None]) -> dict[str, float | int]:
        import statistics
        
        # Filter out None values
        valid_rewards = [r for r in rewards if r is not None]
        
        if not valid_rewards:
            return {}
        
        # Collect all reward keys
        all_keys = set()
        for r in valid_rewards:
            all_keys.update(r.keys())
        
        # Compute median for each reward
        result = {}
        for key in all_keys:
            values = [r[key] for r in valid_rewards if key in r]
            if values:
                result[key] = statistics.median(values)
        
        return result
```

### Metric with Parameters

```python theme={null}
class ThresholdMetric(BaseMetric[dict[str, float]]):
    """Compute percentage of trials above a threshold."""
    
    def __init__(self, threshold: float = 0.8, reward_key: str = "score"):
        self.threshold = threshold
        self.reward_key = reward_key
    
    def compute(self, rewards: list[dict[str, float] | None]) -> dict[str, float | int]:
        valid_rewards = [r for r in rewards if r is not None and self.reward_key in r]
        
        if not valid_rewards:
            return {"above_threshold": 0.0, "n_trials": 0}
        
        above = sum(1 for r in valid_rewards if r[self.reward_key] >= self.threshold)
        percentage = above / len(valid_rewards)
        
        return {
            "above_threshold": percentage,
            "threshold": self.threshold,
            "n_above": above,
            "n_trials": len(valid_rewards)
        }

# Use with custom parameters
metric = ThresholdMetric(threshold=0.9, reward_key="accuracy")
```

### Multi-Reward Analysis Metric

```python theme={null}
class CorrelationMetric(BaseMetric[dict[str, float]]):
    """Compute correlation between two reward metrics."""
    
    def __init__(self, reward1: str, reward2: str):
        self.reward1 = reward1
        self.reward2 = reward2
    
    def compute(self, rewards: list[dict[str, float] | None]) -> dict[str, float | int]:
        import statistics
        
        # Get pairs where both rewards exist
        pairs = []
        for r in rewards:
            if r and self.reward1 in r and self.reward2 in r:
                pairs.append((r[self.reward1], r[self.reward2]))
        
        if len(pairs) < 2:
            return {"correlation": 0.0, "n_pairs": len(pairs)}
        
        # Compute Pearson correlation
        x = [p[0] for p in pairs]
        y = [p[1] for p in pairs]
        
        mean_x = statistics.mean(x)
        mean_y = statistics.mean(y)
        
        numerator = sum((xi - mean_x) * (yi - mean_y) for xi, yi in zip(x, y))
        denominator = (
            sum((xi - mean_x) ** 2 for xi in x) ** 0.5 *
            sum((yi - mean_y) ** 2 for yi in y) ** 0.5
        )
        
        correlation = numerator / denominator if denominator != 0 else 0.0
        
        return {
            "correlation": correlation,
            "n_pairs": len(pairs)
        }
```

## Using Metrics in Jobs

```python theme={null}
from harbor import Job
from harbor.models.job.config import JobConfig
from harbor.models.metric.config import MetricConfig
from harbor.models.metric.type import MetricType

config = JobConfig(
    job_name="comprehensive-metrics",
    agents=[{"name": "claude-code", "model_name": "anthropic/claude-opus-4-1"}],
    datasets=[{"name": "terminal-bench", "version": "2.0"}],
    n_attempts=5,
    metrics=[
        MetricConfig(type=MetricType.MEAN),
        MetricConfig(type=MetricType.MAX),
        MetricConfig(type=MetricType.MIN),
    ]
)

job = Job(config)
result = await job.run()

# Access computed metrics
for eval_key, agent_stats in result.stats.evals.items():
    print(f"\n{eval_key}:")
    for metric in agent_stats.metrics:
        print(f"  {metric}")
```

## Related Types

* [JobConfig](/api/job-config) - Configure metrics for a job
* [JobResult](/api/job-result) - Access computed metrics
* [TrialResult](/api/trial-result) - Individual trial rewards
