Research workflows generate mountains of unstructured data - papers, assay results, genomic annotations, clinical notes - that keyword search cannot adequately surface. Milvus is a purpose-built vector database that indexes high-dimensional embeddings and returns semantically similar results at scale. Synapse, the Sage Bionetwork platform, hosts thousands of biomedical datasets, projects, and files spanning cancer genomics, neuroscience, and rare disease research. Connecting the two lets you embed Synapse content, store those vectors in Milvus, and query them with natural-language prompts - turning a fragmented data lake into a searchable, intelligent research layer your team can query in plain English.
Stream Synapse datasets into Milvus vector collections automatically.
Run similarity searches on enterprise-scale data in milliseconds.
Eliminate manual ETL between your warehouse and vector database.
Semantic Search Across Synapse Projects List projects in Synapse, generate embeddings for their descriptions and metadata, store them in a Milvus collection, and retrieve the top-K most relevant projects for any research query.
File-Level Similarity Search Index individual Synapse files (data dictionaries, README docs, result CSVs) as embedding vectors so researchers can find related files across datasets they may not know exist.
Biomedical Dataset Discovery Browse Synapse biomedical datasets, embed their abstracts and variable definitions, then run similarity searches to find datasets that match a hypothesis or patient cohort profile.
Cross-Project Duplicate and Overlap Detection Search Milvus for near-duplicate embeddings across Synapse project files to identify redundant data collection efforts before a new study begins.
Research Evidence Retrieval Embed published findings stored as Synapse files, index them in Milvus, and retrieve the most semantically relevant evidence when drafting a new grant application or study protocol.
Neotask connects Milvus and Synapse through natural-language instructions. When you describe what you need, Neotask calls Synapse to retrieve the relevant projects or files, passes the content through an embedding model, and upserts the resulting vectors into your Milvus collection with structured metadata. For searches, Neotask embeds your query, runs a similarity search against Milvus, and returns ranked results with direct references back to the originating Synapse entities. No pipeline code is required. You describe the indexing strategy or the search goal, and Neotask orchestrates every API call - from Synapse file listing through Milvus index management - in a single conversation turn.
| Capability | Milvus | Synapse |
|---|---|---|
| List and embed Synapse projects | Store project embeddings | List / get projects |
| Index Synapse files as vectors | Upsert file embeddings | List / get files |
| Similarity search on datasets | Top-K vector query | Browse biomedical datasets |
| Cross-project overlap detection | Near-duplicate search | Multi-project file listing |
| Metadata-filtered vector queries | Filtered ANN search | Project and file metadata |
| Collection management | Create / manage indexes | Dataset enumeration |
Include Synapse entity IDs in Milvus metadata so every search result links directly back to the source project or file - no manual cross-referencing needed.
Use a dedicated Milvus collection per Synapse project when your datasets are large or domain-specific; this keeps index sizes manageable and search precision high.
Re-index on a schedule by asking Neotask to list recently modified Synapse files and upsert only changed vectors - keeping your Milvus collection fresh without full re-indexing.
Neotask orchestrates the indexing based on your instructions, but you control what gets written to Milvus. You can specify which Synapse projects or files to include, and you can delete or update collections at any time by asking Neotask to manage your Milvus indexes. No data is retained by Neotask itself.
Neotask uses the embedding model configured in your environment. You can specify a preferred model in your instructions, such as OpenAI text-embedding-3-small or a locally hosted model. The resulting vectors are stored in the Milvus collection you specify, with dimensionality matching the model output.
Yes. If you index files or project metadata from multiple Synapse projects into the same Milvus collection, a single similarity search query will return ranked results from across all of them. You can also filter by Synapse project ID stored as a metadata field to narrow results to a specific project.
Milvus is designed for billion-scale vector workloads, so the integration scales with your Milvus deployment. For very large Synapse datasets, Neotask can batch the file listing and embedding steps to avoid rate limits and memory pressure, processing records in chunks you define.
You need a Synapse account with at least read access to the projects and files you want to index. For controlled-access biomedical datasets on Synapse, your account must have the appropriate data use approval in place before Neotask can retrieve and embed that content.
Stop hunting through Synapse projects manually. Let Neotask index your research data in Milvus and surface the right datasets, files, and evidence with a single prompt.
$0/mo
Download without a card and start for free.
$50/mo
The full personal agent platform for one person.
$100/mo
One company workspace with room to add your team.
$200/mo
Multiple workspaces and capacity for larger teams.
Explore: Integrations · Skills · Glossary · Solutions · Use cases · Examples · Comparisons · Templates · Blog · Docs