Methodology Paper, Version 1.0 (Draft)
Authored by the SII Initiative · Initial contributor: Alexey Prokofyev · Founder, Arcanis.
Venture capital is a roughly six-hundred-billion-dollar industry whose primary analytical artifact, the investment research report, has no quality standard. Investors are unable to systematically compare the methodologies behind two research products, validate that the same inputs would produce the same conclusions in another analyst’s hands, or trace any specific output to its evidentiary source. The introduction of large language models into research workflows has accelerated production but not raised quality; in many cases it has scaled the existing fragmentation.
This paper proposes Systematic Investment Intelligence (SII), an open framework for measuring the quality of any investment research methodology against five conceptually orthogonal dimensions: Coverage, Method, Proof, Output, and Platform. Each dimension is operationalized through a set of binary, externally testable criteria. The framework is calibrated against an initial sample of nineteen production research platforms in the Growth- and Late-Stage Venture Capital Deep Research category, with empirical results presented. The paper argues that an open standard, analogous in function, though not in scope, to GAAP, GIPS, ILPA, or ISO families, is now both technically feasible and economically necessary, and proposes a governance and certification architecture that prevents any single producer from controlling the standard.
This document is published as a v1.0 draft inviting professional-community review. Subsequent versions will incorporate revisions from independent validators.
Where this fits. SII is an open, methodology-neutral standard, governed independently of any single platform; scores are not endorsements. It is one layer in a broader body of work on private-market pricing, mapped at arcanis.com/research.
The problem. Private-market research lacks a quality standard. Outputs from different research producers are not comparable. Methodologies are typically undocumented. Sources are inconsistently disclosed. There is no equivalent of GAAP, GIPS, or ILPA for the analytical artifact that drives a $400-600B annual deployment decision. The arrival of generative AI has compounded this: AI-produced reports are now indistinguishable in appearance from rigorous research, while typically lacking traceable sourcing, deterministic reproducibility, or auditable reasoning.
The proposal. SII is an open framework that asserts five conceptually orthogonal dimensions of research-methodology quality and, within each dimension, a set of binary criteria with explicit operational tests. The framework is category-agnostic in structure, the five dimensions apply to any analytical workflow, but is category-specific in operationalization. The first category specified is Growth- and Late-Stage VC Deep Research (GLS VC DR); generalizations to Early-stage Venture sourcing, LP allocation research, and public-equity systematic strategy are sketched.
The five dimensions.
Across these five dimensions, twenty-five binary criteria are specified, each with an externally executable test.
Empirical validation. The baseline market scan (nineteen platforms; results presented in Part V) shows that the framework discriminates across the space of current research tools and reveals the structure of the current market. Most platforms cluster in two regimes: structured-system providers (high on Method and Proof simultaneously) and ad-hoc tools (low on both). The empirical correlation between Method and Proof at r = 0.87 is interpreted as a market property, investments in methodology and investments in evidence tend to co-occur in practice, rather than a defect in the framework.
Governance. The framework is published with a proposed governance architecture in which (i) category criteria are submitted by professional contributors, (ii) certification requires three independent third-party validations against a defined protocol, (iii) versions of the standard, the data, and the methodology are tracked together so historical comparisons remain meaningful, and (iv) Alexey Prokofyev, Founder of Arcanis, as initial contributor, holds no veto and no preferential vote.
Status. This is v1.0 of the standard, released as a draft for community review. The current operationalization covers one category (GLS VC DR) and one initial market scan. Open questions and roadmap items for v2.0 are catalogued in Part VIII.
Compared with public markets, private-market research operates in a structurally underdeveloped infrastructure. The public-equity research workflow has converged over four decades into a stack of widely-accepted conventions: GAAP and IFRS for financial reporting; FINRA, MiFID II, and SEC rules for disclosure; CFA-curriculum methodologies for valuation; standardized data feeds; and an audit-and-regulator ecosystem that polices the boundary between assertion and evidence. None of these structural supports exists in private markets at equivalent depth. There is no GAAP for the deal memo. There is no GIPS for the investment thesis. There is no standard structure to a private-company financial model. There is no canonical methodology to value a Series C SaaS company that would compel two analysts at different firms to produce comparable outputs from comparable inputs.
The consequences are concrete and measurable. Survey data consistently shows analyst time-loss to manual data reconciliation in the 30-50% range. Definitions of basic metrics, ARR, NDR, CAC payback, gross retention, vary not only between firms but between memos within the same firm. The reporting standards LPs receive from GPs are unstandardized in format and content. Comparability across funds, vintages, and sectors is a frequently-noted limit on LP decision-making.
For most of the past two decades, this state of affairs was sustainable because the bottleneck on private-market investing was capital, not analysis. Capital was scarce, and the marginal value of better methodology was small relative to the marginal value of better deal access. That equilibrium has flipped. Capital flows have grown faster than analytical capacity. Top-tier firms now have analytical processes that exceed what individual analysts can replicate; the long tail of the market does not. The result is a market structurally bifurcated by infrastructure, with the analytical gap widening rather than closing.
The deployment of large language models into research workflows over 2023-2026 has produced two distinct effects that are commonly conflated.
The first effect is genuine productivity gain. A research task that previously consumed forty hours of analyst time can now be drafted in a few hours with AI assistance, then verified and refined by a human in less time than the original draft would have taken. This is a real and lasting change.
The second effect is the surface-quality illusion. AI-generated research reports look, at the level of formatting and prose, like the output of a senior analyst. They contain section headings, charts, tables, and confident conclusions. They do not, by default, contain the underlying methodology, traceable sources, or deterministic reproducibility that distinguish a rigorous analytical product from a plausible-looking one. The difference is invisible in the deliverable but decisive in the underlying epistemic state of the conclusion.
The combination is a category problem. The cost of producing a research report has fallen by an order of magnitude; the cost of producing a credible research report has not. Without a method to distinguish them, the market is at risk of equilibrating on the lower-cost product, and reaching investment decisions on inputs whose evidentiary backing the decision-maker cannot inspect.
This is the substantive case for a standard. A buyer of research, an LP, a CIO, a family office, a GP committee, needs to be able to ask “does this research meet a defined evidentiary bar?” and receive a meaningful answer. Without that capacity, the market will price all research the same regardless of quality, and producers of higher-quality research will be unable to demonstrate it.
Three precedents are worth holding in mind, with their differences acknowledged.
GAAP / IFRS standardized financial reporting over six decades. They did not solve all problems of comparability, accounting policy choices remain, and earnings management is documented, but they shifted the burden of proof. An investor reading audited financials can assume a structured set of definitions and disclosures unless explicitly told otherwise. The substantive value of GAAP is not the technical detail; it is the existence of a shared scaffolding against which deviations become visible.
GIPS (Global Investment Performance Standards) standardized performance reporting for asset managers, particularly in the institutional discretionary space. GIPS compliance is voluntary; non-compliance is permitted. But the existence of the standard creates a measurable signal: an asset manager who is GIPS-compliant is a different commercial product from one who is not.
ILPA (Institutional Limited Partners Association) templates have standardized LP-GP reporting in private equity in ways that GIPS does not, through structural conventions rather than full performance certification. Adoption is incomplete but meaningful; many LPs now expect ILPA-format quarterly reports as a baseline.
The relevant pattern in each case: a standard does not need to compel adoption to create value. It needs to (a) be conceptually well-formed, (b) be operationally testable, and (c) be governed by an entity that the market regards as neither captured nor merely advisory. The combination of those three properties makes deviation visible. That is the function we propose SII to perform for investment-research methodologies.
SII is presented as a framework with category-specific operationalization. The framework structure is general; the criteria are category-specific.
This v1.0 paper provides: - The framework architecture (Part II). - The empirical orthogonality test and derivation (Part III). - The operationalization for the first defined category, GLS VC Deep Research (Part IV). - Empirical baseline results from a nineteen-platform market scan (Part V). - Sketches of how the same framework would operationalize for other categories (Part VI). - The proposed governance and certification architecture (Part VII). - Open questions and the v2.0 roadmap (Part VIII).
The document does not propose that SII certification be made mandatory. It proposes only that the standard exist and be governable independently of any single producer.
Six principles shape the framework’s structure. They are stated here so that subsequent design choices are inspectable rather than inferred.
P1. Single-home assignment. Every criterion belongs to exactly one dimension. Where v0.x drafts had criteria that could plausibly live in multiple dimensions (e.g., “customization” appearing in three vectors), v1.0 resolves the ambiguity by assigning to the dimension whose definition the criterion most directly tests. The choices are documented and inspectable in the criterion-level notes (Part II, §2.4).
P2. Binary operationalization. Each criterion specifies a test that returns 0 or 1 when executed by an external rater. Tests of quality, completeness, or depth that require subjective judgment are not used; where the underlying property is graded, the test sets a threshold and the binary indicator records whether the threshold is met.
P3. Conceptual orthogonality before empirical orthogonality. The framework is structured so that two criteria in different dimensions are logically independent, one can construct a hypothetical product satisfying one but not the other. This is necessary for the framework to discriminate between products that genuinely differ. The framework does not require empirical orthogonality in the current market sample, because the current market sample is small and structurally bimodal; the framework should still discriminate as the market diversifies.
P4. Category-agnostic structure, category-specific operationalization. The five dimensions describe any research methodology. The criteria within each dimension are specified for a category, currently only GLS VC Deep Research, and may differ across categories. The same framework, with different criteria, would operationalize for Early-stage VC sourcing, LP allocation research, or public-equity systematic strategy.
P5. Version and lineage tracking. Methodology versions, data versions, and result versions are tracked together. A historical result is meaningful only with reference to the standard version under which it was produced.
P6. No producer veto. Alexey Prokofyev, founder of Arcanis, as initial category contributor, holds no governance role over the standard’s evolution beyond the participation rights granted to any other accredited contributor.
The framework asserts that the quality of any investment-research methodology can be characterized along five dimensions, each addressing a distinct question.
| Code | Dimension | Question | Mental model |
|---|---|---|---|
| C | COVERAGE | What is in scope? | “I am not missing anything that matters.” |
| M | METHOD | How is the analysis constructed? | “Same inputs ⇒ same conclusions, by anyone, anywhere, anytime.” |
| P | PROOF | Why should anyone believe the result? | “I can verify every number without trusting the producer.” |
| O | OUTPUT | What does the decision-maker receive? | “Conviction at a glance, drill where I need.” |
| L | PLATFORM | What are the infrastructure properties of the producer? | “It works fast, scales wide, survives five years, and is not owned by an upstream party.” |
These five dimensions correspond directly to the five vectors of v0.9, with renamings introduced where the v0.9 label was internally inconsistent with the dimension’s actual content (e.g., “Complete” in v0.9 contained both input-side and output-side criteria; v1.0 separates them).
The full criterion specification follows. Each criterion includes (a) the question it answers, (b) the binary test used to score it, and (c) where applicable, a note on its lineage from v0.x or its movement between dimensions.
Definition. Coverage is the dimension of analytical input. It asks what the methodology takes as raw material before any processing.
Definition. Method is the dimension of analytical process. It asks how, exactly, the inputs are turned into outputs.
M6 Temporal control. Can the methodology execute as-of any historical date without forward-look leakage? Test: A re-run as of date T uses only data with timestamp ≤ T; the result is identical to one that would have been produced at T.
Definition. Proof is the dimension of evidentiary backing. It asks why a third party should accept the result of the methodology without trusting the producer.
Definition. Output is the dimension of decision utility. It asks what the decision-maker actually receives and whether the form supports the decision.
Definition. Platform is the dimension of infrastructure. It asks about the properties of the system and the organization producing the analysis.
Six substantive changes from the June 2025 framework (v0.9):
These six changes preserve 25 of the 26 v0.9 sub-vectors (with P5 absorbing one) and rearrange six of them into homes more consistent with the dimension definitions.
A framework with five dimensions is orthogonal if knowing a product’s score on one dimension provides no systematic information about its scores on the others. Strict statistical orthogonality (zero correlation across all pairs in any sample) is rarely achievable in human-system frameworks, and it is not necessary for the framework to function. What is necessary is conceptual orthogonality, that the dimensions describe logically independent properties such that one can construct hypothetical products satisfying any subset of dimensions but not others.
Empirical correlation across an observed sample reflects two things: (i) genuine conceptual dependence between dimensions and (ii) population properties of the sample. The two cannot always be separated without a more diverse sample than is currently available. In the v0.9 framework, dimensional correlations were inflated by both effects. In the v1.0 framework, conceptual dependencies have been resolved by single-home assignment; remaining empirical correlations are interpreted as population properties.
The test was conducted on the v0.9 nineteen-platform × twenty-six-criterion scoring matrix, as scored by the SII initial contributor team between April and June 2025. The matrix uses binary 0/1 ratings. The test reports two views:
Sub-criterion view. Pearson correlation between every pair of criteria across the platform sample. Pairs in the same v0.9 dimension are expected to correlate (the criteria describe the same construct from different angles); pairs across dimensions should not. High cross-dimension correlation signals a structural problem.
Dimension view. Correlation between dimension totals (sum of criteria scores per platform) for each pair of dimensions. High cross-dimension correlation indicates either misalignment of criteria or genuine empirical dependence between dimensions in the current market.
The full results are provided in Appendix C. The principal findings:
Three criteria with insufficient discrimination (≥17/19 or ≤2/19 platforms scoring 1): - Universe coverage (17/19 score 1), most platforms claim broad coverage, even when actual depth varies. - Independent research tool (18/19), almost everyone qualifies, suggesting the criterion as worded does not discriminate. - Latest tech stack (18/19), same. - Financial models (2/19), NDA integration (2/19), Backdated research (1/19), these discriminate in favor of one or two platforms and against everyone else.
In v1.0, the near-ceiling criteria are retained because their value is preventive (a 2028 sample will look different if any of these regress) rather than discriminatory in 2025. The near-floor criteria are retained because they identify the high end of the market.
Cross-dimension high correlations (highest in v0.9): - Complete picture at result presentation ↔ Deterministic Repeatable result: r = 1.00, perfect collinearity. The two criteria score identically across all nineteen platforms. The resolution in v1.0 is to move the first to OUTPUT, where it is operationally distinct (O4 section completeness can be assessed independently of M2 reproducibility). - Data and Methodology Versioning ↔ Audit trails: r = 0.89. These remain in different dimensions (METHOD and PROOF) in v1.0; the empirical correlation is interpreted as market property, products that version their methodology also document audit trails. - Zero assumptions ↔ AI providers agnostic: r = 0.88. The Zero-assumptions criterion is merged into Source-grounded evidence in v1.0 (P5), which conceptually distinguishes it from L1 AI-agnosticism.
Dimension-level correlation matrix:
v0.9 (current draft):
Complete S&C T&V Clear F-P Tech Complete 1.00 0.60 0.66 0.36 0.11 S&C 0.60 1.00 0.66 0.31 0.19 T&V 0.66 0.66 1.00 0.32 -0.04 Clear 0.36 0.31 0.32 1.00 0.52 F-P Tech 0.11 0.19 -0.04 0.52 1.00 Mean off-diag |r|: 0.36, max: 0.66
v1.0 (consolidated):
Coverage Method Proof Output Platform Coverage 1.00 0.22 0.28 0.61 0.41 Method 0.22 1.00 0.87 0.53 0.03 Proof 0.28 0.87 1.00 0.49 -0.05 Output 0.61 0.53 0.49 1.00 0.49 Platform 0.41 0.03 -0.05 0.49 1.00 Mean off-diag |r|: 0.40, max: 0.87
The v1.0 mean is slightly higher (0.40 vs 0.36), driven by the METHOD↔PROOF correlation rising to 0.87. We address this directly.
The v1.0 re-homing concentrates two clusters of criteria into METHOD and PROOF that previously straddled three v0.9 dimensions. Empirically they covary at r = 0.87 in the current sample.
Two interpretations are available.
Interpretation A: The dimensions are not actually distinct, and should be merged. Under this view, methodology integrity and evidence transparency are facets of the same underlying construct and the framework should reduce to four dimensions.
Interpretation B: The dimensions are conceptually distinct but empirically co-occurring in the current market. Under this view, the framework should preserve the distinction. Two reasons. First, one can construct products that satisfy METHOD without PROOF (a well-specified methodology that is not documented externally) or PROOF without METHOD (an audit trail of unsystematic reasoning). Second, the current market is structurally bimodal: products either invest in methodology and evidence together, or do neither, but this is a sample property, not a framework property.
We adopt Interpretation B. The METHOD↔PROOF correlation is recorded as a population observation. The framework retains the distinction because it expects intermediate products to emerge as the market diversifies, and because the conceptual distinction matters when defining what failure on each dimension looks like (a failure of METHOD looks like irreproducibility; a failure of PROOF looks like an undocumented but actually-correct procedure).
We note that this design choice is contestable. A v2.0 review may consolidate METHOD and PROOF if the empirical correlation persists in a larger and more diverse sample.
Beyond the correlation structure, three observations about the current market are recorded:
Three reasons. First, growth- and late-stage VC deep research is the analytical workflow with the highest stakes per output: a single research artifact may inform a $10-500M allocation decision. Second, the methodology is mature enough to have canonical components (DCF, comparables, peer benchmarking) while heterogeneous enough that no single producer dominates. Third, the initial contributor (Alexey Prokofyev, founder of Arcanis) has direct production experience in this category, which both enables operational specificity and creates governance risk that the framework must address.
A GLS VC Deep Research artifact is a structured analytical report on a single growth- or late-stage private company, produced for an investment decision. The artifact contains, at minimum: - Company description and business model - Market context and competitive positioning - Financial performance (historical, where available) - Forward-looking projections under a defined scenario set - Risk inventory - Comparable companies (public or private) - Valuation conclusion under at least one canonical model - Recommendation or score relative to a defined investment hypothesis
Scope boundaries: the category includes both single-company deep research and the financial-model artifact alone where it is sold or delivered as the primary product. The category excludes pure data delivery, pure search/discovery tools, and pure CRM/workflow products that do not produce a research artifact.
The scoring protocol for a platform in this category is as follows.
Step 1, Production of a benchmark research artifact. The candidate produces, using the platform under evaluation, a research artifact on a defined benchmark subject. The benchmark subject is selected by the rater from a published shortlist maintained by the SII community.
Step 2, Methodology disclosure submission. The candidate submits the methodology specification used to produce the artifact. The specification must be sufficient for an independent analyst to execute the same protocol.
Step 3, Twenty-five-criterion evaluation. The rater scores each criterion 0 or 1 against the operational test specified for that criterion. Tests requiring computational verification (e.g., M2 deterministic reproducibility) are executed; tests requiring document inspection (e.g., P5 source-grounded evidence) are conducted on a random sample.
Step 4, Result publication. The score is recorded with version stamps for the framework, the methodology specification, and the data corpus. The score is posted to the SII Map (Part VII §7.4) once at least three independent ratings have been completed.
For a binary 25-criterion scoring instrument, the appropriate reliability statistic is Cohen’s kappa per criterion, with Fleiss’s kappa for the three-rater consensus. The protocol target is:
Criteria that fail this threshold in early production rating are flagged for re-specification before being used in published scores. This is the principal mechanism by which the criteria sharpen over time.
For external communication, dimension scores are aggregated into two two-dimensional plots and one one-dimensional indicator. (Original quadrant visualization in v0.9 deck; v1.0 re-anchors the axes to the new dimension labels.)
Plot A: COVERAGE × METHOD. Distinguishes platforms by whether they take in complete inputs (COVERAGE) and process them with integrity (METHOD). Quadrants: Ultimate (high-high), Blindspot (low COVERAGE, high METHOD, well-built but limited in scope), Gambling (high COVERAGE, low METHOD, much data, ad-hoc processing), Fragmented (low-low).
Plot B: PROOF × OUTPUT. Distinguishes platforms by whether their results are verifiable (PROOF) and whether the deliverable is decision-ready (OUTPUT). Quadrants: Science (high-high), Glassmaze (high PROOF, low OUTPUT, verifiable but unusable), Blackbox (low PROOF, high OUTPUT, pretty but unprovable), Darkwater (low-low).
Indicator: PLATFORM. Reported as a single 0-5 score against the platform criteria.
The quadrant labels are retained from v0.9 but reinterpreted against the v1.0 dimensions; the public communication value of the labels (Ultimate / Science / Blindspot / Gambling / Glassmaze / Blackbox / Darkwater / Fragmented) is preserved.
The baseline sample comprises nineteen production platforms operating in or adjacent to the GLS VC DR category as of June 2025. The sample includes:
The sample is not random; it represents the most commercially visible providers in each subcategory at the time of scan. The Arcanis platform is included; its pricing methodology is documented separately in The Decision Surface, and we record the conflict of interest in §5.4.
Scoring was conducted by the Arcanis SII initial contributor team. Each platform was evaluated by hands-on trial where access was available (twelve of nineteen platforms), and by public documentation and demo materials otherwise (seven of nineteen). Each criterion was scored 0 or 1 against the operational test. The full scored matrix is published in Appendix A.
Dimension-level scores (sum of criterion scores within each dimension) for the nineteen platforms, mapped to v1.0:
Platform | Coverage | Method | Proof | Output | Platform |
Arcanis | 4/4 | 6/6 | 5/5 | 5/5 | 4/5 |
OpenAI (DR) | 2/4 | 0/6 | 0/5 | 2/5 | 4/5 |
OpenAI PhD | 3/4 | 0/6 | 0/5 | 3/5 | 4/5 |
Perplexity | 2/4 | 0/6 | 0/5 | 2/5 | 4/5 |
Grok DR | 2/4 | 0/6 | 0/5 | 2/5 | 4/5 |
Gemini | 2/4 | 0/6 | 0/5 | 2/5 | 4/5 |
ScaleX Invest | 3/4 | 3/6 | 3/5 | 5/5 | 4/5 |
Sacra | 0/4 | 2/6 | 2/5 | 2/5 | 1/5 |
Dealpotential | 0/4 | 3/6 | 2/5 | 1/5 | 3/5 |
Mathlabs | 2/4 | 3/6 | 5/5 | 4/5 | 5/5 |
Carried AI | 2/4 | 4/6 | 5/5 | 3/5 | 3/5 |
Velvet | 2/4 | 4/6 | 3/5 | 5/5 | 5/5 |
Auquan | 2/4 | 3/6 | 2/5 | 2/5 | 5/5 |
Dealroom | 1/4 | 2/6 | 2/5 | 1/5 | 1/5 |
CB Insights | 2/4 | 2/6 | 3/5 | 3/5 | 2/5 |
Preqin | 2/4 | 2/6 | 3/5 | 1/5 | 2/5 |
eFront Insights | 2/4 | 2/6 | 3/5 | 1/5 | 3/5 |
AlphaSense | 3/4 | 3/6 | 4/5 | 3/5 | 2/5 |
DealEdge | 2/4 | 2/6 | 3/5 | 2/5 | 2/5 |
(Scores sum to 24/25 maximum because P5 in v1.0 merges two v0.9 criteria.)
The Arcanis platform scores at or near the ceiling on this benchmark scan, and Arcanis is the initial contributor of the framework and the operational tests. This is a material conflict.
Three observations on the conflict.
First, the conflict is structural and unavoidable in a v1.0 release: the initial contributor of any standard has, by construction, designed criteria its own product is built to satisfy. The GIPS, ILPA, and ISO precedents follow the same pattern.
Second, the governance architecture proposed in Part VII addresses this conflict by requiring three independent third-party validations before any certification is recorded, by binding the initial contributor to no voting weight beyond that of any qualified participant, and by versioning the criteria such that retroactive changes are inspectable.
Third, readers of this document should treat the Arcanis row in the baseline scan as a self-assessment, not a certified score. The Arcanis row will be re-scored under the three-rater certification protocol before it is recorded in any public SII Map.
Three observations about the current market are visible in the dimension scores.
A structural divide between platform categories. General LLM products score uniformly high on PLATFORM and low on METHOD and PROOF. Specialized VC platforms cluster in the middle on all five dimensions, with variance. Data providers score high on COVERAGE and moderately on PROOF but low on OUTPUT.
A bimodal distribution on METHOD/PROOF. Most platforms score either close to zero on both (the LLMs and most data providers) or near the ceiling on both (Arcanis, ScaleX, Velvet, Carried AI, Mathlabs). Intermediate scores are rare. This is the empirical pattern that drove the r = 0.87 correlation reported in Part III.
OUTPUT is the dimension where the entire market is weakest relative to its potential. The maximum score on OUTPUT is 5/5, achieved by three platforms (Arcanis, ScaleX, Velvet). Modal score is 2/5. Most platforms produce content but do not produce decision-ready output. This is the dimension most amenable to short-term competitive movement.
The v1.0 specification covers only the GLS VC DR category. The framework architecture is category-agnostic; the criteria are category-specific. This section sketches how the framework would operationalize for three adjacent categories, to demonstrate generality.
In early-stage sourcing, the analytical artifact is not a deep-research report but a screening decision on a high-volume pipeline. COVERAGE shifts from “universe of companies” to “universe of relevant deal flow”; the depth requirement on each subject is lower but the breadth requirement is much higher. METHOD shifts from canonical valuation models (DCF, comparables) to scoring and thematic-fit logic. PROOF still applies, every score must be traceable to its drivers. OUTPUT shifts from decision-readiness on a single subject to ranked or categorized triage across a pipeline. PLATFORM is largely unchanged.
A criterion under M3 (model canonicality) for early-stage might require, for example, “the scoring includes team, market, traction, and thematic-fit factors with declared weights” rather than “the report includes DCF, comparables, and cost.”
In LP allocation research, the subject is not a company but a fund manager or a private-market category. COVERAGE includes manager performance history, vintage exposure, and concentration analysis. METHOD requires standardized comparators (PME, public market equivalents, vintage benchmarks). PROOF is particularly elevated: LPs are accountable to investment committees and trustees who need every claim sourced. OUTPUT includes portfolio-construction analytics rather than single-deal scenarios. PLATFORM is largely unchanged.
Criterion P5 (source-grounded evidence) under LP allocation requires that performance figures cite a verifiable source, typically the GP’s audited reporting or a third-party data feed, and that adjustments are explicit.
In public-equity systematic research, the analytical artifact is a backtest, signal study, or strategy memo. COVERAGE includes the security universe and the historical data corpus. METHOD requires explicit statistical specification, look-ahead avoidance, and out-of-sample testing. PROOF requires that every result is replicable from the disclosed code and data. OUTPUT is the backtest report and the trading signal. PLATFORM is largely unchanged.
The criterion M6 (temporal control) is the binding constraint in this category: any methodology that cannot demonstrate avoidance of look-ahead bias scores 0 regardless of how compelling its results appear.
The five dimensions describe properties of the analytical workflow, not properties of the subject. COVERAGE is about the inputs an analyst takes seriously. METHOD is about the procedure that turns inputs into outputs. PROOF is about why a third party should accept the result. OUTPUT is about what is delivered to the decision-maker. PLATFORM is about the infrastructure of the producer. These five concepts apply to any analytical workflow in finance and, with renaming, to most domains outside finance.
What is category-specific is the operationalization. Future SII categories will define their own criteria within each dimension. The framework architecture is intentionally stable; the criterion library expands.
Three principles structure the proposed governance model.
No producer veto. The initial contributor (Alexey Prokofyev, founder of Arcanis) holds no governance role beyond participation rights granted to any qualified contributor. The standard’s evolution is decided by aggregated contributor votes within rules defined by the SII charter.
Open contribution, qualified participation. Anyone may submit proposed category criteria, proposed methodology improvements, or proposed amendments to the framework. Voting on adoption is restricted to qualified participants, defined in §7.2.
Versioned independence. Every published version of the standard is immutable. Future revisions are released as new versions with documented changes; historical scores remain valid against the version under which they were produced.
A qualified participant is a professional or organization meeting one of the following:
Qualified participants register their AUM (where applicable), role, and category focus. Voting weights on category-criteria amendments are proposed to scale with AUM bands and role categories, subject to a charter constraint that no single participant can exceed 5% of total voting weight on any vote.
The certification levels are:
SII Category Candidate. Self-assessment complete. The candidate has produced a benchmark artifact, submitted methodology documentation, and self-scored against the criterion set. No third-party validation yet.
SII Category Certified. Three independent third-party validations against the published protocol. Each validation produces a score; the certified score is the median across the three rater scores per criterion. Certification is valid for twelve months and must be re-validated annually.
A platform may publish its Candidate-level self-assessment publicly. Certified-level scores are published on the SII Map (§7.4) and may be marketed.
The SII Map is the public record of all Certified scores. It displays, for each category and each platform, the dimension-level and criterion-level scores under the current version of the standard. Historical versions are accessible. Self-assessments (Candidate level) are visible but visually distinguished from Certified scores.
A standard whose initial contributor is also the highest-scoring product is structurally exposed to a producer-capture critique. We acknowledge the exposure and respond on three levels.
First, the criteria themselves are inspectable. Every criterion has a binary operational test. A reader who believes a criterion is gerrymandered toward Arcanis can inspect the test and propose an alternative through the open-contribution process.
Second, the governance constraint is structural: Arcanis’s voting weight is bounded the same way every participant’s is bounded (5% maximum). Arcanis cannot, by charter, block an amendment that the rest of the community supports.
Third, the empirical scan in Part V is presented with the explicit conflict disclosure. The Arcanis row in the published Map will display “Self-assessed pending third-party validation” until and unless the three-rater protocol completes.
These responses do not eliminate the concern. They make it structurally addressable rather than rhetorically dismissible.
Limited category coverage. Only one category (GLS VC DR) is operationalized. Three other category sketches (Part VI) are not yet specified at the criterion level.
Single-rater scoring in baseline. The Part V baseline scan was conducted by the initial contributor team only. The three-rater certification protocol has not yet been executed for any platform.
METHOD/PROOF empirical conflation. The two dimensions correlate at r = 0.87 in the current sample. The framework retains the conceptual distinction (Part III §3.4) but a v2.0 review may revisit this if the empirical pattern persists in a larger sample.
Criterion-set static between major versions. Within a major version, the criterion set is immutable. This is necessary for comparability but limits responsiveness to discovered ambiguity. Minor revisions (clarifications without scoring impact) are permitted; the procedure is specified in the charter.
Inter-rater reliability not yet established. The kappa targets in §4.4 are targets, not measurements. Establishing actual inter-rater reliability requires a multi-rater pilot which is a v1.1 deliverable.
The proposed sequence:
The framework is published under open contribution rules. The four ways to contribute:
The contribution process and the charter governing it are specified in a separate governance document.
| ID | Dimension | Criterion | Binary Test |
|---|---|---|---|
| C1 | Coverage | Universe coverage | ≥90% of a 30-entity test list present |
| C2 | Coverage | Public-source depth | ≥30 sources per major section in corpus |
| C3 | Coverage | Private-data integration | Documented ingestion, access, retention, audit |
| C4 | Coverage | Analytical scope customization | Two users produce materially different scopes |
| M1 | Method | Specified methodology | Independent analyst can execute from spec |
| M2 | Method | Deterministic reproducibility | Two runs, same inputs → identical outputs |
| M3 | Method | Model canonicality | All category-canonical models present |
| M4 | Method | Executor invariance | Swap LLM or analyst, results within tolerance |
| M5 | Method | Method and data versioning | Every result tagged with method and data version |
| M6 | Method | Temporal control | As-of date T uses only ≤T data; result is what would have been produced at T |
| P1 | Proof | Methodology disclosure | Every rule findable in user-facing docs |
| P2 | Proof | Logic-tree audit trail | Drill from output to source through every step |
| P3 | Proof | Reproducible calculation | External party regenerates any number |
| P4 | Proof | AI explainability | Every AI-produced sentence traceable to sources/reasoning |
| P5 | Proof | Source-grounded evidence | 20/20 random claims have working source citations |
| O1 | Output | Single-view decision readiness | “Invest/pass/dig” from one screen, ≥80% of subjects |
| O2 | Output | Scenario navigation | Side-by-side scenarios with delta per metric |
| O3 | Output | Attention-optimized structure | Decision-relevant data within first 60s |
| O4 | Output | Section completeness | All canonical sections present, ordered |
| O5 | Output | Scenario universe interaction | Input change propagates without re-run |
| L1 | Platform | AI-provider agnosticism | Two LLM providers, results within tolerance |
| L2 | Platform | Modern tech stack | All major deps within active vendor support |
| L3 | Platform | Non-regression test automation | CI blocks release on drift beyond tolerance |
| L4 | Platform | Producer independence | Disclosed commercial relationships; no conflicts in scoring logic |
| L5 | Platform | Delivery latency | Median end-to-end ≤4h (VC DR) |
The framework has three documented predecessor drafts:
v0.7, URBAN (early 2025). Four dimensions (Ultimate Instrument, Complete Data, Systematic Approach, Transparent) with 23 sub-features. Focused on Arcanis’s product differentiation rather than as an open standard.
v0.8, Expanded SII (mid 2025). Five dimensions with 33 sub-features, including sub-vectors that v0.9 dropped: Complete scenarios’ universe, Most complete & evolving methods, Methodology customization (separate from Modular framework), Rapid adaptation of newest research methods, Interactive output dashboard (separate from scenario navigation), Automated-Attention-Optimized Output, Scalable and localizable architecture.
v0.9, Consolidated SII (June 2025 deck). Five dimensions with 26 sub-features. The version reproduced in the SII community research deck. Dimensions: Complete, Systematic & Custom, Transparent & Verifiable, Clear Result, Future-proof Tech.
v1.0, Orthogonality-tested SII (this document). Five dimensions with 25 criteria. Restructuring relative to v0.9 documented in §2.4 and §3.3.
The v0.8 sub-features that were dropped in v0.9 and not restored in v1.0 are recorded here for completeness:
Scalable and localizable architecture, partially covered by L2 (modern stack) and L5 (delivery latency); the v0.8 version was not operationally testable.
SII, Systematic Investment Intelligence. The framework defined in this document.
Category, A defined class of investment-research workflow (e.g., GLS VC DR). Each category has its own criterion specification within the general framework.
Criterion, A binary, externally-testable property of a research methodology.
Dimension, One of the five top-level properties (Coverage, Method, Proof, Output, Platform).
Operational test, The specific procedure by which a rater determines whether a criterion is satisfied.
Qualified participant, An individual or organization meeting the participation criteria in §7.2.
SII Map, The public record of Certified platform scores under the framework.
Version, A released specification of the framework or of a category criterion set. Major versions (v1, v2) are immutable; minor versions add clarifications without changing scoring.
End of v1.0 draft. Comments and amendments to: [governance address to be assigned by SII charter].