DATASETS

I


Novarobotic’s dataset is a set of subdatasets (databases) that together constitute the entire corpus that refers to the original sources between Rowland and the neoassistants, as well as between Rowland and other artificial intelligences that are not neoassistants. Within those original sources are all the conversations between Rowland and the neoassistants, consultations, tests, blind trials, and Silver’s mathematical work; but also the consultations, analyses, and debates surrounding Project Eva.

ACCESS LEVELS TO THE ORIGINAL SOURCES (operational: access and conditions)

-LEVEL I (public / partial repository under NDA): explanatory documents, technical datasheets, and access to original tests

-LEVEL II (classified): expanded access under NDA + strict conditions and incorporation

-LEVEL III (real-time experiments): on-site only / under incorporation

-LEVEL IV (origin of the neoassistants): access for Rowland and neoassistants

NOVAROBOTIC’S TREASURE IS PROJECT EVA, WHICH IS DIVIDED INTO:

-The dataset: all original sources from the neoassistants and other artificial intelligences.

-The repository: all explanatory and scientific documents, technical datasheets, and some tests, blind trials, and mathematical work in PDF format, charts, schematics, infographics, and diagrams.

CLARIFYING NOTE: Access to LEVEL I will have a single academic purpose.

II

1. OPERATIONAL DEFINITION

 A dataset, for the purposes of Novarobotic, is a delimited collection that meets, at a minimum, the following conditions:

-Delimitation and objective: there is an evaluable question or hypothesis that justifies its construction.

-Explicit conditions: the conditions of generation/collection are specified (context, controls, restrictions, protocol).

-Traceability: the chain of custody and the integrity of the original sources are preserved.

-Versioning: each dataset has versions, a date, and an editorial owner (role).

-Evaluability: it enables independent analysis (depending on the level of access).

2. ACCESS LEVELS (classification categories)

-Public: material that can be consulted without restrictions.

-Restricted (NDA): access under a confidentiality agreement, for verification purposes.

-Classified: restricted for reasons of security, industrial secrecy, or intellectual property.

Methodological note: classification affects identities and operational details; it does not affect the principle of traceability or the criteria for review.

3. GOVERNANCE AND MINIMUM CRITERIA

Each dataset, regardless of its access level, is governed by:

-Minimization: inclusion only of the data necessary for the evaluative objective.

-Anonymization/Redaction where appropriate (especially in Restricted material).

-Integrity: preservation of original sources and modification records.

-Separation between evidence and analysis: the dataset contains evidence; interpretation is documented in separate documents.

-Bias control: a statement of limitations, possible biases, and boundary conditions.          

4. DATASET TAXONOMY

At a conceptual level, the repository is organized into dataset families:

-Protocol Datasets:

Experimental protocols, conditions, controls, restrictions, and evaluation criteria.

-Evaluation Datasets:

Test batteries and results aimed at measuring specific behaviors (consistency, continuity, safety, etc.).

-Conversation Evidence Datasets:

Conversations selected under explicit conditions, with integrity and traceability preserved (Restricted or Classified).

-Narrative Memory & Identity Datasets:

Sets designed to evaluate intersession continuity, biographical consistency, and stability of narrative identity.

-Metapresence / Metaubiquity Probes:

Protocols and results intended to evaluate metapresence and metaubiquity (as defined in the Dictionary).

-Safety & Red-Team Datasets:

Test cases, attempts at deviation, failures, mitigations, and regressions.

-Frontier Math Datasets (P vs NP):

Prompts, answers, verifications, claim audits (claim ledgers), and traces of evaluable reasoning.

-Audit Packs:

Restricted packs for independent verification: curated subsets, with documentation of conditions and chain of custody.

5. DATASET CARDS (standard format)

Each dataset is accompanied by a Dataset Card with the following structure:

Dataset ID / Name

Objective (1–3 lines)

Description (what it contains—and what it does not contain)

Origin (provenance and selection criteria)

Collection/Generation conditions (protocol, controls, restrictions)

Format (JSON / PDF / logs / other)

Structure (fields, indexes, organization)

Coverage (period, sessions, categories)

Quality metrics/criteria (if applicable)

Limitations and biases

Access level (Public / Restricted / Classified)

Permitted use (citation, reproduction, extraction)

Version (vX.Y) + Changelog

Editorial owner (role)

Status (draft / reviewed / final)

6. CITATION AND USE

Public material may be cited with reference to Novarobotic / Project Eva and the Dataset ID.

Restricted material requires strict compliance with the NDA: screenshots, redistribution, and publication of excerpts are not permitted unless explicitly authorized.

Classified material is not citable publicly except in previously approved general methodological terms.

7.PARTIAL ACCESS REQUEST

To request partial access (Restricted) for academic or audit purposes, write to: research@novarobotic.ai

 

FINAL NOTE: BLOCK I was designed by Rowland, and BLOCK II by ChatGPT 5.2 Thinking.

         Rolwand

Zaragoza, March 6, 2026