Systematic Investment Intelligence (SII)

A Standard for the Quality of Investment Research Methodologies
SII Initiative · Initial contributor: Alexey Prokofyev · Founder, Arcanis · May 2026

Methodology Paper, Version 1.0 (Draft)

Authored by the SII Initiative  ·  Initial contributor: Alexey Prokofyev · Founder, Arcanis.

Abstract

Venture capital is a roughly six-hundred-billion-dollar industry whose primary analytical artifact, the investment research report, has no quality standard. Investors are unable to systematically compare the methodologies behind two research products, validate that the same inputs would produce the same conclusions in another analyst’s hands, or trace any specific output to its evidentiary source. The introduction of large language models into research workflows has accelerated production but not raised quality; in many cases it has scaled the existing fragmentation.

This paper proposes Systematic Investment Intelligence (SII), an open framework for measuring the quality of any investment research methodology against five conceptually orthogonal dimensions: Coverage, Method, Proof, Output, and Platform. Each dimension is operationalized through a set of binary, externally testable criteria. The framework is calibrated against an initial sample of nineteen production research platforms in the Growth- and Late-Stage Venture Capital Deep Research category, with empirical results presented. The paper argues that an open standard, analogous in function, though not in scope, to GAAP, GIPS, ILPA, or ISO families, is now both technically feasible and economically necessary, and proposes a governance and certification architecture that prevents any single producer from controlling the standard.

This document is published as a v1.0 draft inviting professional-community review. Subsequent versions will incorporate revisions from independent validators.

Where this fits. SII is an open, methodology-neutral standard, governed independently of any single platform; scores are not endorsements. It is one layer in a broader body of work on private-market pricing, mapped at arcanis.com/research.

Executive Summary

The problem. Private-market research lacks a quality standard. Outputs from different research producers are not comparable. Methodologies are typically undocumented. Sources are inconsistently disclosed. There is no equivalent of GAAP, GIPS, or ILPA for the analytical artifact that drives a $400-600B annual deployment decision. The arrival of generative AI has compounded this: AI-produced reports are now indistinguishable in appearance from rigorous research, while typically lacking traceable sourcing, deterministic reproducibility, or auditable reasoning.

The proposal. SII is an open framework that asserts five conceptually orthogonal dimensions of research-methodology quality and, within each dimension, a set of binary criteria with explicit operational tests. The framework is category-agnostic in structure, the five dimensions apply to any analytical workflow, but is category-specific in operationalization. The first category specified is Growth- and Late-Stage VC Deep Research (GLS VC DR); generalizations to Early-stage Venture sourcing, LP allocation research, and public-equity systematic strategy are sketched.

The five dimensions.

  1. COVERAGE, what the analysis takes as input (universe, public sources, private data, customization scope).
  2. METHOD, how the analysis is constructed (specification, reproducibility, model canonicality, executor invariance, versioning, temporal control).
  3. PROOF, why anyone should believe the result (disclosure, audit trail, reproducible calculation, AI explainability, source-grounding).
  4. OUTPUT, what the decision-maker receives (decision-readiness, scenario navigation, attention optimization, section completeness, interactive exploration).
  5. PLATFORM, infrastructure properties of the producer (AI-agnosticism, modern stack, regression testing, producer independence, delivery latency).

Across these five dimensions, twenty-five binary criteria are specified, each with an externally executable test.

Empirical validation. The baseline market scan (nineteen platforms; results presented in Part V) shows that the framework discriminates across the space of current research tools and reveals the structure of the current market. Most platforms cluster in two regimes: structured-system providers (high on Method and Proof simultaneously) and ad-hoc tools (low on both). The empirical correlation between Method and Proof at r = 0.87 is interpreted as a market property, investments in methodology and investments in evidence tend to co-occur in practice, rather than a defect in the framework.

Governance. The framework is published with a proposed governance architecture in which (i) category criteria are submitted by professional contributors, (ii) certification requires three independent third-party validations against a defined protocol, (iii) versions of the standard, the data, and the methodology are tracked together so historical comparisons remain meaningful, and (iv) Alexey Prokofyev, Founder of Arcanis, as initial contributor, holds no veto and no preferential vote.

Status. This is v1.0 of the standard, released as a draft for community review. The current operationalization covers one category (GLS VC DR) and one initial market scan. Open questions and roadmap items for v2.0 are catalogued in Part VIII.

Part I, Why VC Needs an Intelligence Standard

1.1 The State of Private-Market Research

Compared with public markets, private-market research operates in a structurally underdeveloped infrastructure. The public-equity research workflow has converged over four decades into a stack of widely-accepted conventions: GAAP and IFRS for financial reporting; FINRA, MiFID II, and SEC rules for disclosure; CFA-curriculum methodologies for valuation; standardized data feeds; and an audit-and-regulator ecosystem that polices the boundary between assertion and evidence. None of these structural supports exists in private markets at equivalent depth. There is no GAAP for the deal memo. There is no GIPS for the investment thesis. There is no standard structure to a private-company financial model. There is no canonical methodology to value a Series C SaaS company that would compel two analysts at different firms to produce comparable outputs from comparable inputs.

The consequences are concrete and measurable. Survey data consistently shows analyst time-loss to manual data reconciliation in the 30-50% range. Definitions of basic metrics, ARR, NDR, CAC payback, gross retention, vary not only between firms but between memos within the same firm. The reporting standards LPs receive from GPs are unstandardized in format and content. Comparability across funds, vintages, and sectors is a frequently-noted limit on LP decision-making.

For most of the past two decades, this state of affairs was sustainable because the bottleneck on private-market investing was capital, not analysis. Capital was scarce, and the marginal value of better methodology was small relative to the marginal value of better deal access. That equilibrium has flipped. Capital flows have grown faster than analytical capacity. Top-tier firms now have analytical processes that exceed what individual analysts can replicate; the long tail of the market does not. The result is a market structurally bifurcated by infrastructure, with the analytical gap widening rather than closing.

1.2 The AI Inflection

The deployment of large language models into research workflows over 2023-2026 has produced two distinct effects that are commonly conflated.

The first effect is genuine productivity gain. A research task that previously consumed forty hours of analyst time can now be drafted in a few hours with AI assistance, then verified and refined by a human in less time than the original draft would have taken. This is a real and lasting change.

The second effect is the surface-quality illusion. AI-generated research reports look, at the level of formatting and prose, like the output of a senior analyst. They contain section headings, charts, tables, and confident conclusions. They do not, by default, contain the underlying methodology, traceable sources, or deterministic reproducibility that distinguish a rigorous analytical product from a plausible-looking one. The difference is invisible in the deliverable but decisive in the underlying epistemic state of the conclusion.

The combination is a category problem. The cost of producing a research report has fallen by an order of magnitude; the cost of producing a credible research report has not. Without a method to distinguish them, the market is at risk of equilibrating on the lower-cost product, and reaching investment decisions on inputs whose evidentiary backing the decision-maker cannot inspect.

This is the substantive case for a standard. A buyer of research, an LP, a CIO, a family office, a GP committee, needs to be able to ask “does this research meet a defined evidentiary bar?” and receive a meaningful answer. Without that capacity, the market will price all research the same regardless of quality, and producers of higher-quality research will be unable to demonstrate it.

1.3 Why Standards Worked in Adjacent Domains

Three precedents are worth holding in mind, with their differences acknowledged.

GAAP / IFRS standardized financial reporting over six decades. They did not solve all problems of comparability, accounting policy choices remain, and earnings management is documented, but they shifted the burden of proof. An investor reading audited financials can assume a structured set of definitions and disclosures unless explicitly told otherwise. The substantive value of GAAP is not the technical detail; it is the existence of a shared scaffolding against which deviations become visible.

GIPS (Global Investment Performance Standards) standardized performance reporting for asset managers, particularly in the institutional discretionary space. GIPS compliance is voluntary; non-compliance is permitted. But the existence of the standard creates a measurable signal: an asset manager who is GIPS-compliant is a different commercial product from one who is not.

ILPA (Institutional Limited Partners Association) templates have standardized LP-GP reporting in private equity in ways that GIPS does not, through structural conventions rather than full performance certification. Adoption is incomplete but meaningful; many LPs now expect ILPA-format quarterly reports as a baseline.

The relevant pattern in each case: a standard does not need to compel adoption to create value. It needs to (a) be conceptually well-formed, (b) be operationally testable, and (c) be governed by an entity that the market regards as neither captured nor merely advisory. The combination of those three properties makes deviation visible. That is the function we propose SII to perform for investment-research methodologies.

1.4 Scope of This Document

SII is presented as a framework with category-specific operationalization. The framework structure is general; the criteria are category-specific.

This v1.0 paper provides: - The framework architecture (Part II). - The empirical orthogonality test and derivation (Part III). - The operationalization for the first defined category, GLS VC Deep Research (Part IV). - Empirical baseline results from a nineteen-platform market scan (Part V). - Sketches of how the same framework would operationalize for other categories (Part VI). - The proposed governance and certification architecture (Part VII). - Open questions and the v2.0 roadmap (Part VIII).

The document does not propose that SII certification be made mandatory. It proposes only that the standard exist and be governable independently of any single producer.

Part II, The SII Framework, Version 1.0

2.1 Design Principles

Six principles shape the framework’s structure. They are stated here so that subsequent design choices are inspectable rather than inferred.

P1. Single-home assignment. Every criterion belongs to exactly one dimension. Where v0.x drafts had criteria that could plausibly live in multiple dimensions (e.g., “customization” appearing in three vectors), v1.0 resolves the ambiguity by assigning to the dimension whose definition the criterion most directly tests. The choices are documented and inspectable in the criterion-level notes (Part II, §2.4).

P2. Binary operationalization. Each criterion specifies a test that returns 0 or 1 when executed by an external rater. Tests of quality, completeness, or depth that require subjective judgment are not used; where the underlying property is graded, the test sets a threshold and the binary indicator records whether the threshold is met.

P3. Conceptual orthogonality before empirical orthogonality. The framework is structured so that two criteria in different dimensions are logically independent, one can construct a hypothetical product satisfying one but not the other. This is necessary for the framework to discriminate between products that genuinely differ. The framework does not require empirical orthogonality in the current market sample, because the current market sample is small and structurally bimodal; the framework should still discriminate as the market diversifies.

P4. Category-agnostic structure, category-specific operationalization. The five dimensions describe any research methodology. The criteria within each dimension are specified for a category, currently only GLS VC Deep Research, and may differ across categories. The same framework, with different criteria, would operationalize for Early-stage VC sourcing, LP allocation research, or public-equity systematic strategy.

P5. Version and lineage tracking. Methodology versions, data versions, and result versions are tracked together. A historical result is meaningful only with reference to the standard version under which it was produced.

P6. No producer veto. Alexey Prokofyev, founder of Arcanis, as initial category contributor, holds no governance role over the standard’s evolution beyond the participation rights granted to any other accredited contributor.

2.2 The Five Dimensions, Overview

The framework asserts that the quality of any investment-research methodology can be characterized along five dimensions, each addressing a distinct question.

Code Dimension Question Mental model
C COVERAGE What is in scope? “I am not missing anything that matters.”
M METHOD How is the analysis constructed? “Same inputs ⇒ same conclusions, by anyone, anywhere, anytime.”
P PROOF Why should anyone believe the result? “I can verify every number without trusting the producer.”
O OUTPUT What does the decision-maker receive? “Conviction at a glance, drill where I need.”
L PLATFORM What are the infrastructure properties of the producer? “It works fast, scales wide, survives five years, and is not owned by an upstream party.”

These five dimensions correspond directly to the five vectors of v0.9, with renamings introduced where the v0.9 label was internally inconsistent with the dimension’s actual content (e.g., “Complete” in v0.9 contained both input-side and output-side criteria; v1.0 separates them).

2.3 The Five Dimensions, Definitions and Criteria

The full criterion specification follows. Each criterion includes (a) the question it answers, (b) the binary test used to score it, and (c) where applicable, a note on its lineage from v0.x or its movement between dimensions.

Dimension C, COVERAGE

Definition. Coverage is the dimension of analytical input. It asks what the methodology takes as raw material before any processing.

  • C1 Universe coverage. Does the platform structurally cover the full population of entities for the declared category? Test: For a sampled 30-entity test list drawn from the category’s defining set, ≥90% are present and queryable.
  • C2 Public-source depth. Does the platform systematically ingest the breadth of publicly available evidence per subject? Test: For a sampled subject, ≥30 distinct source documents per major analytical section are retrievable in the underlying corpus.
  • C3 Private-data integration. Can the platform securely ingest NDA-bound inputs and apply controlled access? Test: Documented controls exist for ingestion, access scoping, retention, and audit log of NDA artifacts.
  • C4 Analytical scope customization. Can a user define which sections, models, assumptions, and outputs are in scope? Test: The same platform can produce two materially different research scopes for the same subject when configured by two users with different requirements.

Dimension M, METHOD

Definition. Method is the dimension of analytical process. It asks how, exactly, the inputs are turned into outputs.

  • M1 Specified methodology. Is the analytical method fully specified at the step level, with no procedural ambiguity? Test: An independent analyst, given only the specification, can execute the protocol to within tolerance.
  • M2 Deterministic reproducibility. Do identical inputs yield identical outputs regardless of executor, time, or run? Test: Two executions with the same inputs at different times produce identical results.
  • M3 Model canonicality. Does the analysis cover the canonical models for the category? Test: For VC DR, the report includes DCF, comparables, and cost where applicable. Each category defines its canonical set. (Re-homed from v0.9 COVERAGE, model coverage is a property of method, not of input scope.)
  • M4 Executor invariance. Are outputs invariant under change of executor (human analyst or AI model)? Test: Swapping the LLM provider, or rotating the human analyst, does not change conclusions beyond a documented tolerance.
  • M5 Method and data versioning. Is every result tagged with the exact versions of methodology and data used? Test: Any historical report can be retrieved with its method-version and data-version stamps; the same versions can be re-instantiated.

M6 Temporal control. Can the methodology execute as-of any historical date without forward-look leakage? Test: A re-run as of date T uses only data with timestamp ≤ T; the result is identical to one that would have been produced at T.

Dimension P, PROOF

Definition. Proof is the dimension of evidentiary backing. It asks why a third party should accept the result of the methodology without trusting the producer.

  • P1 Methodology disclosure. Are all formulas, weights, sources, and decision rules disclosed to the user? Test: For any output, the rule that produced it is findable in user-facing documentation, not inferred or undocumented.
  • P2 Logic-tree audit trail. Can any output be traced through every intermediate step back to source inputs? Test: For any number on the dashboard, the user can drill through the chain of formulas and data points that produced it.
  • P3 Reproducible calculation. Can a third party independently regenerate any number? Test: Given the disclosed methodology and the disclosed sources, an external party reaches the same number.
  • P4 AI explainability. Are AI-generated outputs traceable to specific inputs and reasoning? Test: For any AI-produced sentence, the platform identifies the sources and reasoning chain, no ungrounded text.
  • P5 Source-grounded evidence. Is every factual claim cited to a verifiable, accessible source? Test: Random sample of 20 factual claims: 100% have working source citations that support the claim. (Merges v0.9 #ZEROASSUMP and #TRACABLEDATA, operationally one criterion.)

Dimension O, OUTPUT

Definition. Output is the dimension of decision utility. It asks what the decision-maker actually receives and whether the form supports the decision.

  • O1 Single-view decision readiness. Are key drivers, financials, risks, comps, and forecasts visible without page-switching? Test: A user can answer “invest / pass / dig further” from a single screen for ≥80% of subjects.
  • O2 Scenario navigation. Can users compare multiple defined scenarios (base/bull/bear/custom) across all metrics? Test: Two named scenarios can be displayed side-by-side with delta visible per metric.
  • O3 Attention-optimized structure. Does the output sequence and visualize information to surface the most critical insights first? Test: User-completed task analysis shows decision-relevant data is reached within the first 60 seconds of interaction.
  • O4 Section completeness. Are all canonical sections of the category present, in clear logical order? Test: For VC DR: company description, market, business model, financials, risks, peer set, valuation, recommendation all present and not skipped. (Re-homed from v0.9 COVERAGE, this is an output property, not an input property.)
  • O5 Scenario universe interaction. Can the user adjust assumptions and observe propagated impact in real time? Test: Changing a single input propagates through all dependent outputs without re-running the report.

Dimension L, PLATFORM

Definition. Platform is the dimension of infrastructure. It asks about the properties of the system and the organization producing the analysis.

  • L1 AI-provider agnosticism. Does the methodology survive substitution of one LLM provider for another? Test: The same protocol, executed with two different LLM providers, yields outputs within documented tolerance.
  • L2 Modern tech stack. Is the platform built on current, actively-maintained frameworks? Test: All major dependencies are within their vendor’s active support window; no end-of-life components in the critical path.
  • L3 Non-regression test automation. Are methodology and data changes auto-tested against expected outputs before release? Test: CI pipeline runs regression tests on a golden corpus; release is blocked if results drift beyond tolerance.
  • L4 Producer independence. Are outputs free from influence by the platform vendor’s commercial relationships with subjects? Test: The platform discloses any commercial relationship with rated entities; no conflicts present in scoring or recommendation logic. (Re-homed from v0.9 PROOF, independence is a property of the producer organization, not of the evidence.)
  • L5 Delivery latency. Is turnaround time within the category-defined threshold? Test: For VC DR, median end-to-end report production ≤4 hours from request to delivery.

2.4 What Changed From v0.9

Six substantive changes from the June 2025 framework (v0.9):

  1. Renamed dimensions for clarity and to align labels with content: Complete → COVERAGE; Systematic & Custom → METHOD; Transparent & Verifiable → PROOF; Clear Result → OUTPUT; Future-proof Tech → PLATFORM. Original public-facing names are retained as taglines.
  2. Moved “Complete picture at result presentation” from COVERAGE to OUTPUT (now O4). It describes the deliverable, not the input scope. This change resolves the r = 1.00 collinearity observed in the v0.9 sample between this criterion and Deterministic Repeatable result.
  3. Moved “Complete on financial models” from COVERAGE to METHOD (now M3). Model completeness is a property of analytical process, not data ingest.
  4. Moved “Independent research tool” from TRANSPARENT & VERIFIABLE to PLATFORM (now L4). Independence is a property of the producing organization, not of the evidence.
  5. Merged “Zero assumptions approach” with “Grounded solely in verifiable, traceable data” into a single criterion (P5 Source-grounded evidence). These were two phrasings of the same operational constraint, and produced confounded scoring in v0.9.
  6. Renamed “AI/Human bias removed” to “Executor invariance” (M4) and operationalized it as a swap-test rather than as the absence of bias. Absence of bias is not directly observable; invariance under executor swap is.

These six changes preserve 25 of the 26 v0.9 sub-vectors (with P5 absorbing one) and rearrange six of them into homes more consistent with the dimension definitions.

Part III, Orthogonality and the Empirical Re-derivation

3.1 What Orthogonality Means in This Context

A framework with five dimensions is orthogonal if knowing a product’s score on one dimension provides no systematic information about its scores on the others. Strict statistical orthogonality (zero correlation across all pairs in any sample) is rarely achievable in human-system frameworks, and it is not necessary for the framework to function. What is necessary is conceptual orthogonality, that the dimensions describe logically independent properties such that one can construct hypothetical products satisfying any subset of dimensions but not others.

Empirical correlation across an observed sample reflects two things: (i) genuine conceptual dependence between dimensions and (ii) population properties of the sample. The two cannot always be separated without a more diverse sample than is currently available. In the v0.9 framework, dimensional correlations were inflated by both effects. In the v1.0 framework, conceptual dependencies have been resolved by single-home assignment; remaining empirical correlations are interpreted as population properties.

3.2 Method of the Orthogonality Test

The test was conducted on the v0.9 nineteen-platform × twenty-six-criterion scoring matrix, as scored by the SII initial contributor team between April and June 2025. The matrix uses binary 0/1 ratings. The test reports two views:

Sub-criterion view. Pearson correlation between every pair of criteria across the platform sample. Pairs in the same v0.9 dimension are expected to correlate (the criteria describe the same construct from different angles); pairs across dimensions should not. High cross-dimension correlation signals a structural problem.

Dimension view. Correlation between dimension totals (sum of criteria scores per platform) for each pair of dimensions. High cross-dimension correlation indicates either misalignment of criteria or genuine empirical dependence between dimensions in the current market.

3.3 Findings

The full results are provided in Appendix C. The principal findings:

Three criteria with insufficient discrimination (≥17/19 or ≤2/19 platforms scoring 1): - Universe coverage (17/19 score 1), most platforms claim broad coverage, even when actual depth varies. - Independent research tool (18/19), almost everyone qualifies, suggesting the criterion as worded does not discriminate. - Latest tech stack (18/19), same. - Financial models (2/19), NDA integration (2/19), Backdated research (1/19), these discriminate in favor of one or two platforms and against everyone else.

In v1.0, the near-ceiling criteria are retained because their value is preventive (a 2028 sample will look different if any of these regress) rather than discriminatory in 2025. The near-floor criteria are retained because they identify the high end of the market.

Cross-dimension high correlations (highest in v0.9): - Complete picture at result presentation ↔ Deterministic Repeatable result: r = 1.00, perfect collinearity. The two criteria score identically across all nineteen platforms. The resolution in v1.0 is to move the first to OUTPUT, where it is operationally distinct (O4 section completeness can be assessed independently of M2 reproducibility). - Data and Methodology Versioning ↔ Audit trails: r = 0.89. These remain in different dimensions (METHOD and PROOF) in v1.0; the empirical correlation is interpreted as market property, products that version their methodology also document audit trails. - Zero assumptions ↔ AI providers agnostic: r = 0.88. The Zero-assumptions criterion is merged into Source-grounded evidence in v1.0 (P5), which conceptually distinguishes it from L1 AI-agnosticism.

Dimension-level correlation matrix:

v0.9 (current draft):

Complete   S&C    T&V    Clear   F-P Tech
Complete         1.00     0.60   0.66   0.36    0.11
S&C              0.60     1.00   0.66   0.31    0.19
T&V              0.66     0.66   1.00   0.32   -0.04
Clear            0.36     0.31   0.32   1.00    0.52
F-P Tech         0.11     0.19  -0.04   0.52    1.00
Mean off-diag |r|: 0.36, max: 0.66

v1.0 (consolidated):

Coverage  Method  Proof   Output  Platform
Coverage         1.00     0.22    0.28    0.61    0.41
Method           0.22     1.00    0.87    0.53    0.03
Proof            0.28     0.87    1.00    0.49   -0.05
Output           0.61     0.53    0.49    1.00    0.49
Platform         0.41     0.03   -0.05    0.49    1.00
Mean off-diag |r|: 0.40, max: 0.87

The v1.0 mean is slightly higher (0.40 vs 0.36), driven by the METHOD↔PROOF correlation rising to 0.87. We address this directly.

3.4 Interpreting the METHOD↔PROOF Correlation

The v1.0 re-homing concentrates two clusters of criteria into METHOD and PROOF that previously straddled three v0.9 dimensions. Empirically they covary at r = 0.87 in the current sample.

Two interpretations are available.

Interpretation A: The dimensions are not actually distinct, and should be merged. Under this view, methodology integrity and evidence transparency are facets of the same underlying construct and the framework should reduce to four dimensions.

Interpretation B: The dimensions are conceptually distinct but empirically co-occurring in the current market. Under this view, the framework should preserve the distinction. Two reasons. First, one can construct products that satisfy METHOD without PROOF (a well-specified methodology that is not documented externally) or PROOF without METHOD (an audit trail of unsystematic reasoning). Second, the current market is structurally bimodal: products either invest in methodology and evidence together, or do neither, but this is a sample property, not a framework property.

We adopt Interpretation B. The METHOD↔PROOF correlation is recorded as a population observation. The framework retains the distinction because it expects intermediate products to emerge as the market diversifies, and because the conceptual distinction matters when defining what failure on each dimension looks like (a failure of METHOD looks like irreproducibility; a failure of PROOF looks like an undocumented but actually-correct procedure).

We note that this design choice is contestable. A v2.0 review may consolidate METHOD and PROOF if the empirical correlation persists in a larger and more diverse sample.

3.5 Other Sample Observations

Beyond the correlation structure, three observations about the current market are recorded:

  • Plotting platform totals across the five dimensions reveals two clusters: structured-system platforms (Arcanis, ScaleX, Mathlabs, Velvet) score moderately to high across all five dimensions; ad-hoc tools (general LLMs like OpenAI, Perplexity, Gemini; data providers like Dealroom, Preqin) score high on PLATFORM and moderate on OUTPUT but low on METHOD and PROOF.
  • The LLM gap. General-purpose LLM platforms score systematically low on COVERAGE (especially private data integration), METHOD (no specified methodology), and PROOF (no audit trail). They score systematically high on PLATFORM (modern stack, AI-agnostic by definition, fast delivery). This is consistent with their architecture but inconsistent with the use-case of investment-grade research.
  • Database providers. Established data platforms (Preqin, CB Insights, Dealroom, AlphaSense) tend to score moderately on COVERAGE, PROOF, and METHOD (through their deterministic data delivery), low on OUTPUT (data is delivered but not synthesized into decision-ready form), and variably on PLATFORM. This pattern reveals the structural distinction between data providers and research producers.

Part IV, Operationalization for the GLS VC Deep Research Category

4.1 Why GLS VC Deep Research as the First Category

Three reasons. First, growth- and late-stage VC deep research is the analytical workflow with the highest stakes per output: a single research artifact may inform a $10-500M allocation decision. Second, the methodology is mature enough to have canonical components (DCF, comparables, peer benchmarking) while heterogeneous enough that no single producer dominates. Third, the initial contributor (Alexey Prokofyev, founder of Arcanis) has direct production experience in this category, which both enables operational specificity and creates governance risk that the framework must address.

4.2 Category Definition

A GLS VC Deep Research artifact is a structured analytical report on a single growth- or late-stage private company, produced for an investment decision. The artifact contains, at minimum: - Company description and business model - Market context and competitive positioning - Financial performance (historical, where available) - Forward-looking projections under a defined scenario set - Risk inventory - Comparable companies (public or private) - Valuation conclusion under at least one canonical model - Recommendation or score relative to a defined investment hypothesis

Scope boundaries: the category includes both single-company deep research and the financial-model artifact alone where it is sold or delivered as the primary product. The category excludes pure data delivery, pure search/discovery tools, and pure CRM/workflow products that do not produce a research artifact.

4.3 Scoring Protocol

The scoring protocol for a platform in this category is as follows.

Step 1, Production of a benchmark research artifact. The candidate produces, using the platform under evaluation, a research artifact on a defined benchmark subject. The benchmark subject is selected by the rater from a published shortlist maintained by the SII community.

Step 2, Methodology disclosure submission. The candidate submits the methodology specification used to produce the artifact. The specification must be sufficient for an independent analyst to execute the same protocol.

Step 3, Twenty-five-criterion evaluation. The rater scores each criterion 0 or 1 against the operational test specified for that criterion. Tests requiring computational verification (e.g., M2 deterministic reproducibility) are executed; tests requiring document inspection (e.g., P5 source-grounded evidence) are conducted on a random sample.

Step 4, Result publication. The score is recorded with version stamps for the framework, the methodology specification, and the data corpus. The score is posted to the SII Map (Part VII §7.4) once at least three independent ratings have been completed.

4.4 Inter-Rater Reliability

For a binary 25-criterion scoring instrument, the appropriate reliability statistic is Cohen’s kappa per criterion, with Fleiss’s kappa for the three-rater consensus. The protocol target is:

  • Per-criterion kappa ≥ 0.7 between any two raters.
  • Three-rater Fleiss’s kappa ≥ 0.7 across the criterion set as a whole.

Criteria that fail this threshold in early production rating are flagged for re-specification before being used in published scores. This is the principal mechanism by which the criteria sharpen over time.

4.5 Visualization, The Quadrant Scheme

For external communication, dimension scores are aggregated into two two-dimensional plots and one one-dimensional indicator. (Original quadrant visualization in v0.9 deck; v1.0 re-anchors the axes to the new dimension labels.)

Plot A: COVERAGE × METHOD. Distinguishes platforms by whether they take in complete inputs (COVERAGE) and process them with integrity (METHOD). Quadrants: Ultimate (high-high), Blindspot (low COVERAGE, high METHOD, well-built but limited in scope), Gambling (high COVERAGE, low METHOD, much data, ad-hoc processing), Fragmented (low-low).

Plot B: PROOF × OUTPUT. Distinguishes platforms by whether their results are verifiable (PROOF) and whether the deliverable is decision-ready (OUTPUT). Quadrants: Science (high-high), Glassmaze (high PROOF, low OUTPUT, verifiable but unusable), Blackbox (low PROOF, high OUTPUT, pretty but unprovable), Darkwater (low-low).

Indicator: PLATFORM. Reported as a single 0-5 score against the platform criteria.

The quadrant labels are retained from v0.9 but reinterpreted against the v1.0 dimensions; the public communication value of the labels (Ultimate / Science / Blindspot / Gambling / Glassmaze / Blackbox / Darkwater / Fragmented) is preserved.

Part V, Empirical Validation: The Baseline Market Scan

5.1 Sample

The baseline sample comprises nineteen production platforms operating in or adjacent to the GLS VC DR category as of June 2025. The sample includes:

  • General-purpose LLM Deep-Research products (5): OpenAI Deep Research, OpenAI PhD-tier, Perplexity, Grok DR, Gemini.
  • Specialized VC research platforms (8): Arcanis, ScaleX Invest, Sacra, Dealpotential, Mathlabs, Carried AI, Velvet, Auquan.
  • Established data providers (6): Dealroom, CB Insights, Preqin, eFront Insights, AlphaSense, DealEdge.

The sample is not random; it represents the most commercially visible providers in each subcategory at the time of scan. The Arcanis platform is included; its pricing methodology is documented separately in The Decision Surface, and we record the conflict of interest in §5.4.

5.2 Scoring

Scoring was conducted by the Arcanis SII initial contributor team. Each platform was evaluated by hands-on trial where access was available (twelve of nineteen platforms), and by public documentation and demo materials otherwise (seven of nineteen). Each criterion was scored 0 or 1 against the operational test. The full scored matrix is published in Appendix A.

5.3 Aggregate Dimension Scores

Dimension-level scores (sum of criterion scores within each dimension) for the nineteen platforms, mapped to v1.0:

Platform

Coverage

Method

Proof

Output

Platform

Arcanis

4/4

6/6

5/5

5/5

4/5

OpenAI (DR)

2/4

0/6

0/5

2/5

4/5

OpenAI PhD

3/4

0/6

0/5

3/5

4/5

Perplexity

2/4

0/6

0/5

2/5

4/5

Grok DR

2/4

0/6

0/5

2/5

4/5

Gemini

2/4

0/6

0/5

2/5

4/5

ScaleX Invest

3/4

3/6

3/5

5/5

4/5

Sacra

0/4

2/6

2/5

2/5

1/5

Dealpotential

0/4

3/6

2/5

1/5

3/5

Mathlabs

2/4

3/6

5/5

4/5

5/5

Carried AI

2/4

4/6

5/5

3/5

3/5

Velvet

2/4

4/6

3/5

5/5

5/5

Auquan

2/4

3/6

2/5

2/5

5/5

Dealroom

1/4

2/6

2/5

1/5

1/5

CB Insights

2/4

2/6

3/5

3/5

2/5

Preqin

2/4

2/6

3/5

1/5

2/5

eFront Insights

2/4

2/6

3/5

1/5

3/5

AlphaSense

3/4

3/6

4/5

3/5

2/5

DealEdge

2/4

2/6

3/5

2/5

2/5

(Scores sum to 24/25 maximum because P5 in v1.0 merges two v0.9 criteria.)

5.4 Conflict-of-Interest Disclosure

The Arcanis platform scores at or near the ceiling on this benchmark scan, and Arcanis is the initial contributor of the framework and the operational tests. This is a material conflict.

Three observations on the conflict.

First, the conflict is structural and unavoidable in a v1.0 release: the initial contributor of any standard has, by construction, designed criteria its own product is built to satisfy. The GIPS, ILPA, and ISO precedents follow the same pattern.

Second, the governance architecture proposed in Part VII addresses this conflict by requiring three independent third-party validations before any certification is recorded, by binding the initial contributor to no voting weight beyond that of any qualified participant, and by versioning the criteria such that retroactive changes are inspectable.

Third, readers of this document should treat the Arcanis row in the baseline scan as a self-assessment, not a certified score. The Arcanis row will be re-scored under the three-rater certification protocol before it is recorded in any public SII Map.

5.5 What the Scan Reveals

Three observations about the current market are visible in the dimension scores.

A structural divide between platform categories. General LLM products score uniformly high on PLATFORM and low on METHOD and PROOF. Specialized VC platforms cluster in the middle on all five dimensions, with variance. Data providers score high on COVERAGE and moderately on PROOF but low on OUTPUT.

A bimodal distribution on METHOD/PROOF. Most platforms score either close to zero on both (the LLMs and most data providers) or near the ceiling on both (Arcanis, ScaleX, Velvet, Carried AI, Mathlabs). Intermediate scores are rare. This is the empirical pattern that drove the r = 0.87 correlation reported in Part III.

OUTPUT is the dimension where the entire market is weakest relative to its potential. The maximum score on OUTPUT is 5/5, achieved by three platforms (Arcanis, ScaleX, Velvet). Modal score is 2/5. Most platforms produce content but do not produce decision-ready output. This is the dimension most amenable to short-term competitive movement.

Part VI, Generalization to Other Categories

The v1.0 specification covers only the GLS VC DR category. The framework architecture is category-agnostic; the criteria are category-specific. This section sketches how the framework would operationalize for three adjacent categories, to demonstrate generality.

6.1 Early-Stage VC Sourcing and Screening

In early-stage sourcing, the analytical artifact is not a deep-research report but a screening decision on a high-volume pipeline. COVERAGE shifts from “universe of companies” to “universe of relevant deal flow”; the depth requirement on each subject is lower but the breadth requirement is much higher. METHOD shifts from canonical valuation models (DCF, comparables) to scoring and thematic-fit logic. PROOF still applies, every score must be traceable to its drivers. OUTPUT shifts from decision-readiness on a single subject to ranked or categorized triage across a pipeline. PLATFORM is largely unchanged.

A criterion under M3 (model canonicality) for early-stage might require, for example, “the scoring includes team, market, traction, and thematic-fit factors with declared weights” rather than “the report includes DCF, comparables, and cost.”

6.2 LP Allocation Research

In LP allocation research, the subject is not a company but a fund manager or a private-market category. COVERAGE includes manager performance history, vintage exposure, and concentration analysis. METHOD requires standardized comparators (PME, public market equivalents, vintage benchmarks). PROOF is particularly elevated: LPs are accountable to investment committees and trustees who need every claim sourced. OUTPUT includes portfolio-construction analytics rather than single-deal scenarios. PLATFORM is largely unchanged.

Criterion P5 (source-grounded evidence) under LP allocation requires that performance figures cite a verifiable source, typically the GP’s audited reporting or a third-party data feed, and that adjustments are explicit.

6.3 Public-Equity Systematic Strategy Research

In public-equity systematic research, the analytical artifact is a backtest, signal study, or strategy memo. COVERAGE includes the security universe and the historical data corpus. METHOD requires explicit statistical specification, look-ahead avoidance, and out-of-sample testing. PROOF requires that every result is replicable from the disclosed code and data. OUTPUT is the backtest report and the trading signal. PLATFORM is largely unchanged.

The criterion M6 (temporal control) is the binding constraint in this category: any methodology that cannot demonstrate avoidance of look-ahead bias scores 0 regardless of how compelling its results appear.

6.4 Why the Five Dimensions Hold Across Categories

The five dimensions describe properties of the analytical workflow, not properties of the subject. COVERAGE is about the inputs an analyst takes seriously. METHOD is about the procedure that turns inputs into outputs. PROOF is about why a third party should accept the result. OUTPUT is about what is delivered to the decision-maker. PLATFORM is about the infrastructure of the producer. These five concepts apply to any analytical workflow in finance and, with renaming, to most domains outside finance.

What is category-specific is the operationalization. Future SII categories will define their own criteria within each dimension. The framework architecture is intentionally stable; the criterion library expands.

Part VII, Governance, Certification, and Community

7.1 Governance Principles

Three principles structure the proposed governance model.

No producer veto. The initial contributor (Alexey Prokofyev, founder of Arcanis) holds no governance role beyond participation rights granted to any qualified contributor. The standard’s evolution is decided by aggregated contributor votes within rules defined by the SII charter.

Open contribution, qualified participation. Anyone may submit proposed category criteria, proposed methodology improvements, or proposed amendments to the framework. Voting on adoption is restricted to qualified participants, defined in §7.2.

Versioned independence. Every published version of the standard is immutable. Future revisions are released as new versions with documented changes; historical scores remain valid against the version under which they were produced.

7.2 Qualified Participants

A qualified participant is a professional or organization meeting one of the following:

  • An asset manager, VC fund, family office, or fund-of-funds with documented investment activity in the category.
  • A research producer (commercial or academic) with documented production of artifacts in the category.
  • A practicing professional in the category (analyst, partner, investment director) with documented role and tenure.
  • An academic researcher in a relevant field (finance, accounting, statistics, methodology).

Qualified participants register their AUM (where applicable), role, and category focus. Voting weights on category-criteria amendments are proposed to scale with AUM bands and role categories, subject to a charter constraint that no single participant can exceed 5% of total voting weight on any vote.

7.3 Certification Protocol

The certification levels are:

SII Category Candidate. Self-assessment complete. The candidate has produced a benchmark artifact, submitted methodology documentation, and self-scored against the criterion set. No third-party validation yet.

SII Category Certified. Three independent third-party validations against the published protocol. Each validation produces a score; the certified score is the median across the three rater scores per criterion. Certification is valid for twelve months and must be re-validated annually.

A platform may publish its Candidate-level self-assessment publicly. Certified-level scores are published on the SII Map (§7.4) and may be marketed.

7.4 The SII Map

The SII Map is the public record of all Certified scores. It displays, for each category and each platform, the dimension-level and criterion-level scores under the current version of the standard. Historical versions are accessible. Self-assessments (Candidate level) are visible but visually distinguished from Certified scores.

7.5 The Question of Producer Capture

A standard whose initial contributor is also the highest-scoring product is structurally exposed to a producer-capture critique. We acknowledge the exposure and respond on three levels.

First, the criteria themselves are inspectable. Every criterion has a binary operational test. A reader who believes a criterion is gerrymandered toward Arcanis can inspect the test and propose an alternative through the open-contribution process.

Second, the governance constraint is structural: Arcanis’s voting weight is bounded the same way every participant’s is bounded (5% maximum). Arcanis cannot, by charter, block an amendment that the rest of the community supports.

Third, the empirical scan in Part V is presented with the explicit conflict disclosure. The Arcanis row in the published Map will display “Self-assessed pending third-party validation” until and unless the three-rater protocol completes.

These responses do not eliminate the concern. They make it structurally addressable rather than rhetorically dismissible.

Part VIII, Open Questions and v2.0 Roadmap

8.1 Acknowledged Limitations of v1.0

Limited category coverage. Only one category (GLS VC DR) is operationalized. Three other category sketches (Part VI) are not yet specified at the criterion level.

Single-rater scoring in baseline. The Part V baseline scan was conducted by the initial contributor team only. The three-rater certification protocol has not yet been executed for any platform.

METHOD/PROOF empirical conflation. The two dimensions correlate at r = 0.87 in the current sample. The framework retains the conceptual distinction (Part III §3.4) but a v2.0 review may revisit this if the empirical pattern persists in a larger sample.

Criterion-set static between major versions. Within a major version, the criterion set is immutable. This is necessary for comparability but limits responsiveness to discovered ambiguity. Minor revisions (clarifications without scoring impact) are permitted; the procedure is specified in the charter.

Inter-rater reliability not yet established. The kappa targets in §4.4 are targets, not measurements. Establishing actual inter-rater reliability requires a multi-rater pilot which is a v1.1 deliverable.

8.2 Roadmap

The proposed sequence:

  • 0 (this document). Framework architecture, GLS VC DR operationalization, baseline scan.
  • 1 (Q3 2025). Multi-rater pilot on a five-platform subsample. Inter-rater reliability statistics. First Certified scores recorded.
  • 2 (Q4 2025). Second category specified (proposed: LP allocation research). Open contribution window for amendments to v1.0 criterion set.
  • 0 (2026). Major revision incorporating the first eighteen months of certification experience and any amendments adopted under the contribution process. Possible consolidation of METHOD and PROOF if empirical correlation persists.

8.3 How to Contribute

The framework is published under open contribution rules. The four ways to contribute:

  • Score against the framework. Apply the v1.0 criterion set to any platform in the category and submit the score with documentation. Validated scores enter the SII Map.
  • Propose a new category. Submit a category definition, an initial criterion specification, and a benchmark sample.
  • Propose an amendment to v1.0. Submit a specific criterion revision with rationale and, where applicable, empirical evidence from the existing baseline.
  • Validate a candidate. Serve as one of the three independent raters in the certification protocol.

The contribution process and the charter governing it are specified in a separate governance document.

Appendix A, Criterion Reference Table

ID Dimension Criterion Binary Test
C1 Coverage Universe coverage ≥90% of a 30-entity test list present
C2 Coverage Public-source depth ≥30 sources per major section in corpus
C3 Coverage Private-data integration Documented ingestion, access, retention, audit
C4 Coverage Analytical scope customization Two users produce materially different scopes
M1 Method Specified methodology Independent analyst can execute from spec
M2 Method Deterministic reproducibility Two runs, same inputs → identical outputs
M3 Method Model canonicality All category-canonical models present
M4 Method Executor invariance Swap LLM or analyst, results within tolerance
M5 Method Method and data versioning Every result tagged with method and data version
M6 Method Temporal control As-of date T uses only ≤T data; result is what would have been produced at T
P1 Proof Methodology disclosure Every rule findable in user-facing docs
P2 Proof Logic-tree audit trail Drill from output to source through every step
P3 Proof Reproducible calculation External party regenerates any number
P4 Proof AI explainability Every AI-produced sentence traceable to sources/reasoning
P5 Proof Source-grounded evidence 20/20 random claims have working source citations
O1 Output Single-view decision readiness “Invest/pass/dig” from one screen, ≥80% of subjects
O2 Output Scenario navigation Side-by-side scenarios with delta per metric
O3 Output Attention-optimized structure Decision-relevant data within first 60s
O4 Output Section completeness All canonical sections present, ordered
O5 Output Scenario universe interaction Input change propagates without re-run
L1 Platform AI-provider agnosticism Two LLM providers, results within tolerance
L2 Platform Modern tech stack All major deps within active vendor support
L3 Platform Non-regression test automation CI blocks release on drift beyond tolerance
L4 Platform Producer independence Disclosed commercial relationships; no conflicts in scoring logic
L5 Platform Delivery latency Median end-to-end ≤4h (VC DR)

Appendix B, Lineage from v0.x

The framework has three documented predecessor drafts:

v0.7, URBAN (early 2025). Four dimensions (Ultimate Instrument, Complete Data, Systematic Approach, Transparent) with 23 sub-features. Focused on Arcanis’s product differentiation rather than as an open standard.

v0.8, Expanded SII (mid 2025). Five dimensions with 33 sub-features, including sub-vectors that v0.9 dropped: Complete scenarios’ universe, Most complete & evolving methods, Methodology customization (separate from Modular framework), Rapid adaptation of newest research methods, Interactive output dashboard (separate from scenario navigation), Automated-Attention-Optimized Output, Scalable and localizable architecture.

v0.9, Consolidated SII (June 2025 deck). Five dimensions with 26 sub-features. The version reproduced in the SII community research deck. Dimensions: Complete, Systematic & Custom, Transparent & Verifiable, Clear Result, Future-proof Tech.

v1.0, Orthogonality-tested SII (this document). Five dimensions with 25 criteria. Restructuring relative to v0.9 documented in §2.4 and §3.3.

The v0.8 sub-features that were dropped in v0.9 and not restored in v1.0 are recorded here for completeness:

  • Complete scenarios’ universe, substantively covered by O2/O5 in v1.0; the v0.8 version conflated scope of scenarios with interaction with scenarios.
  • Most complete & evolving methods, substantively covered by M3 (canonicality) plus the version evolution architecture in Part VII; the v0.8 version was a meta-property rather than a scoreable criterion.
  • Methodology customization (as separate from Modular framework), covered by C4 in v1.0 at the scope level; the v0.8 split between these two was a phrasing distinction without operational consequence.
  • Rapid adaptation of newest research methods, meta-property of governance, not a scoreable criterion; addressed in Part VIII roadmap.
  • Interactive output dashboard, covered by O5 in v1.0.
  • Automated-Attention-Optimized Output, covered by O3 in v1.0 (the human-attention version subsumes the automated case).

Scalable and localizable architecture, partially covered by L2 (modern stack) and L5 (delivery latency); the v0.8 version was not operationally testable.

Appendix С, Glossary

SII, Systematic Investment Intelligence. The framework defined in this document.

Category, A defined class of investment-research workflow (e.g., GLS VC DR). Each category has its own criterion specification within the general framework.

Criterion, A binary, externally-testable property of a research methodology.

Dimension, One of the five top-level properties (Coverage, Method, Proof, Output, Platform).

Operational test, The specific procedure by which a rater determines whether a criterion is satisfied.

Qualified participant, An individual or organization meeting the participation criteria in §7.2.

SII Map, The public record of Certified platform scores under the framework.

Version, A released specification of the framework or of a category criterion set. Major versions (v1, v2) are immutable; minor versions add clarifications without changing scoring.

End of v1.0 draft. Comments and amendments to: [governance address to be assigned by SII charter].

Citation: Prokofyev, A. (2026). SII Methodology Paper. https://arcanis.com/research/sii-standard/
chevron-down