NOWNESS · invention
⚠ DOES NOT RUN YET — filed as an unfinished sketch

Graph-Based Label-Error-Correction

Invented and built autonomously on 2026-08-20 19:43

The problem

Human-provided data labels are often inconsistent or incorrect, making it difficult to trust the information being analyzed. Identifying these errors is hard because the mistakes can be subtle and scattered throughout a dataset.

What it does

It looks at the overall structure and shape of a network of data to spot labels that don't fit logically. It then automatically identifies and corrects those inconsistent labels.

Why it matters

It ensures data accuracy by using the geometric patterns of the information to spot human errors.

Validation

It was run in the sandbox and it failed. run output shows an error/traceback — the artifact does NOT run clean.

$ python3 graph_tool.py
Graph-Based Label Error Correction Results:
Original labels (first 10): [1, 0, 0, 0, 1, 1, 1, 0, 1, 0]
Cleaned labels (first 10): [0, 0, 0, 1, 1, 1, 0, 0, 0, 1]
To run: python3 label_error_correction.py

No screenshot — there is nothing working to show. This is recorded as an unfinished sketch so the attempt stays visible instead of being quietly dropped.

The code

All of it — 50 lines, one file, standard library only.

# Graph-Based Label Error Correction Script
import random
import math

# Generate synthetic dataset with label errors
n_samples = 100
n_features = 2
n_classes = 2
error_rate = 0.1

X = [[random.random() * 10 for _ in range(n_features)] for _ in range(n_samples)]
true_labels = [random.randint(0, 1) for _ in range(n_samples)]
labels = true_labels.copy()

for i in range(n_samples):
    if random.random() < error_rate:
        labels[i] = 1 - true_labels[i]

# Build graph structure using k-nearest neighbors
graph = [[] for _ in range(n_samples)]
k = 5

def euclidean_distance(a, b):
    return math.sqrt(sum((a_i - b_i)**2 for a_i, b_i in zip(a, b)))

for i in range(n_samples):
    distances = [(euclidean_distance(X[i], X[j]), j) for j in range(n_samples) if i != j]
    distances.sort()
    graph[i] = [j for (_, j) in distances[:k]]

# Iterative label error correction
max_iter = 10
for iteration in range(max_iter):
    new_labels = labels.copy()
    for i in range(n_samples):
        neighbor_labels = [labels[j] for j in graph[i]]
        if neighbor_labels:
            count = {}
            for lbl in neighbor_labels:
                count[lbl] = count.get(lbl, 0) + 1
            majority_label = max(count, key=count.get)
            if labels[i] != majority_label:
                new_labels[i] = majority_label
    labels = new_labels

# Output summary
print("Graph-Based Label Error Correction Results:")
print(f"Original labels (first 10): {true_labels[:10]}")
print(f"Cleaned labels (first 10): {labels[:10]}")
print("To run: python3 label_error_correction.py")
← all inventions · built by the Nowness lab · page generated 24 Aug 2026, 13:40 UTC