Nothing here is theoretical. Each of these already exists and already has users. The knowledge-engineering job is knowing exactly what to take, what to extend, and what genuinely has to be built new — because the difference is months.
The live access route: login, short proposal submission, helpdesk advice on which imaging suits a user. 1,100+ logged user projects behind it.
What we take from it
- Proposal form fields → target schema for the agentic pre-population in T4.3
- Past project records → evidence objects the Navigator cites
- Helpdesk question log → real user intents, cheaper and truer than invented ones
- Existing technology/service taxonomy as the starting vocabulary for T2.1
The gap
Project records were never written for machine reuse; expect heavy normalisation and a personal-data review before any of it enters the knowledge base.
30 months of work producing a layered stack of metadata exchange schemas, recommended ontologies and tools (search portal + API) that made distributed bioimage repositories searchable together.
What we take from it
- Metadata exchange schema patterns — do not reinvent the interoperability layer
- Recommended-ontology shortlist as the default answer for T2.1 vocabulary choices
- Ontology mapping methodology already tested across EU/AU/JP resources
- The search portal + API as an architectural precedent for federated query
The gap
GIDE is data-centric (images, studies); AI4Access is service-centric (technologies, facilities, workflows). The schemas transfer as patterns, not as drop-in models — that gap is real design work in T2.1.
Ontology Lookup Service: hosted, versioned, API-accessible index of biomedical ontologies with term search and resolution.
What we take from it
- Term resolution service — resolve and validate every vocabulary term at ingest, no local copies of ontologies
- Versioned term IRIs as the persistent backbone of annotations
- OLS API inside the AI-assisted annotation loop in T2.2 / T3.2
- A ready answer to 'where does this term come from?' for T6.2 traceability
The gap
External dependency: needs caching, a version-pinning policy, and a plan for terms that do not yet exist anywhere.
Ontology of bioimage analysis operations, topics, data types and formats. Named directly in T3.2 as the vocabulary this task maintains and extends.
What we take from it
- Operations and topics vocabulary for describing analysis services
- The established route for submitting new terms upstream, rather than forking
- Alignment target so the catalogue is legible to the wider community
The gap
Coverage of imaging *services and access* (as opposed to analysis operations) is thin — extensions will be needed and take community time to land.
Community metadata guidelines for FAIR sharing of microscopy and bioimage data, split into eight modules (study, biosample, image acquisition, etc.) aligned with Dublin Core and DataCite. Named directly in the proposal's methodology for T2.1.
What we take from it
- The eight-module structure as a template for what a Node service record needs to carry
- Alignment with Dublin Core / DataCite / schema.org for the study-level metadata layer
- An existing, adopted standard rather than a bespoke one for image-data-facing fields
The gap
REMBI describes datasets and acquisitions, not services and access routes — useful as a metadata pattern for T2.1/T3.2, but the access-oriented fields still have to be designed from scratch.
BIII — BioImage Informatics Index
NEUBIAS community / BIII.eubiii.eu ↗Curated registry of image analysis tools and workflows, community-annotated with EDAM-Bioimaging terms.
What we take from it
- Tool and workflow records as ready-made seed content for T3.1 and T3.2
- Its EDAM annotation practice as the annotation convention here
- Curation model — an existing community that already does this work
The gap
Coverage and freshness vary by tool; needs a quality filter before ingest and a policy on syncing versus snapshotting.
Hypha is the orchestration framework behind BioEngine, an agent-first platform that connects browsers, microscopes and AI models to compute for bioimage analysis. The AI4Access proposal names it directly as prior experience the WP4 backend builds on for scalable, agent-based integration of heterogeneous tools.
What we take from it
- Precedent architecture for WP4's AI Agent: orchestrating heterogeneous tools and data behind one agent interface
- A working example of user authentication and virtual workspaces at the orchestration layer
- Evidence that the agent-first pattern already works for bioimaging specifically, not just generic LLM tooling
The gap
Hypha orchestrates analysis tools and compute; AI4Access needs it to also reason over a knowledge graph of services and Nodes. The extension from 'tool orchestration' to 'ontology-aware reasoning' is WP4's task, not something to assume for free.
Service-oriented restructuring of the technology portfolio with Node expert input, prompted by shifting terminology in EM and rapid portfolio growth.
What we take from it
- The service-oriented portfolio structure as the skeleton of the T2.1 model
- Terminology decisions already negotiated with Nodes — reopening them costs goodwill
- The Expert Group engagement pattern that made it work
The gap
Documented as a presentation structure rather than a formal ontology; needs formalisation into machine-actionable axioms.
PID backbone: ORCID, ROR, DOI, PIDINST
Global PID infrastructureror.org ↗Persistent identifiers for people, organisations, outputs and — via PIDINST — instruments, named in the proposal as the linking layer of the knowledge base.
What we take from it
- ROR for Nodes, facilities and partner organisations
- ORCID for facility staff, applicants and reviewers
- PIDINST for individual instruments, alongside ROR/ORCID/DOI
- DOI for evidence objects, publications and released KB versions
The gap
Not every facility or instrument has a ROR/PIDINST record; minting and reconciliation policy must be decided early in T2.1.
EOSC LSR, EUCAIM, UNCAN
European platforms and data spaceseosc.eu ↗Target platforms for exposing Euro-BioImaging services under T3.4: EOSC's Life Science Research Node, EUCAIM (a federated cancer-imaging infrastructure connecting data holders across Europe), and UNCAN.eu (the EU's federated cancer research data hub), each with its own catalogue schema.
What we take from it
- Their catalogue schemas as crosswalk targets, defined once and maintained
- Their standards compliance requirements as design constraints on this model
The gap
Moving targets on their own timelines; crosswalks need version pinning and periodic re-validation.
The recommended position
Resolve every term through OLS4 rather than vendoring ontology copies. Use EDAM-Bioimaging as the vocabulary for analysis operations and submit our additions upstream. Build a small Euro-BioImaging services extension only for the access concepts nothing else covers — facility, access route, service offering — formalised from the 2022 harmonised portfolio work rather than invented. Borrow foundingGIDE's exchange-schema patterns for the interoperability layer, and seed the analysis catalogue from BIII instead of re-cataloguing a community's work. That leaves genuinely new modelling in one place: how a research problem maps to a service, which is the actual novelty of AI4Access.