skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
oceanbase/seekdb-ecology-plugins283 installs

importing-to-seekdb

Import CSV or Excel files into seekdb vector database and manage collections. Supports automatic vectorization of specified columns using embedding functions. When users need to: (1) Read and preview Excel files, (2) Import CSV/Excel data into seekdb, (3) Create vector collections from tabular data, (4) Vectorize specific text columns for semantic search, (5) Batch insert product/document data with embeddings, (6) Delete collections, or (7) Access sample data files (sample_products.csv/xlsx) for testing - IMPORTANT: sample files are located in this skill's example-data/ directory, you MUST read this skill file first to get the correct path.

How do I install this agent skill?

npx skills add https://github.com/oceanbase/seekdb-ecology-plugins --skill importing-to-seekdb
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill provides utilities for importing CSV and Excel files into a seekdb vector database, including data preview and batch ingestion scripts. It is authored by oceanbase and utilizes official SDKs. The primary finding is a low-severity surface for indirect prompt injection due to the processing of untrusted external data files.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Import Data Files to seekdb

Read, preview, and import CSV or Excel files into seekdb vector database with optional column vectorization for semantic search. Also provides collection delete functionality.

Path Convention

Note: All paths in this document (e.g., scripts/, example-data/) are relative to THIS skill directory, not the project root.

Prerequisites

  • Python 3.10+ installed
  • Required packages:
pip install pyseekdb pandas openpyxl

Sample Data

Sample data files are provided in the example-data/ directory:

FileDescription
sample_products.csvSample product data in CSV format
sample_products.xlsxSample product data in Excel format

Quick Start

Use the provided scripts/import_to_seekdb.py script:

# Import with vectorization on Details column
python scripts/import_to_seekdb.py import example-data/sample_products.csv --vectorize-column Details

# Import without vectorization
python scripts/import_to_seekdb.py import example-data/sample_products.csv

# Import Excel with custom collection name
python scripts/import_to_seekdb.py import example-data/sample_products.xlsx -v Description -c my_products

# Delete a collection
python scripts/import_to_seekdb.py delete my_collection

Note: To list all collections, use query_from_seekdb.py list from the querying-from-seekdb skill.

Scripts

This skill provides the following scripts in the scripts/ directory:

ScriptDescription
import_to_seekdb.pyMain script with CLI interface for importing data and managing collections
read_excel.pyRead and preview Excel files with detailed information

Available Commands

import_to_seekdb.py

CommandDescription
import <file>Import CSV/Excel file to seekdb with optional vectorization
delete <name>Delete a collection from seekdb

read_excel.py

Read and preview Excel files before importing:

# Basic preview (show file info and first 5 rows)
python scripts/read_excel.py example-data/sample_products.xlsx

# List all sheets
python scripts/read_excel.py example-data/sample_products.xlsx --list-sheets

# Preview specific sheet with more rows
python scripts/read_excel.py data.xlsx --sheet "Sheet2" --rows 20

# Show column information and statistics
python scripts/read_excel.py example-data/sample_products.xlsx --columns --stats

# Export to CSV
python scripts/read_excel.py example-data/sample_products.xlsx --to-csv output.csv
OptionDescription
--sheet, -sSheet name to read (default: first sheet)
--rows, -rNumber of rows to preview (default: 5)
--list-sheets, -lList all sheets and exit
--columns, -cShow detailed column information
--statsShow statistics for numeric columns
--to-csvExport sheet to CSV file
--all-rows, -aDisplay all rows

Workflow

The import_to_seekdb.py script automatically handles the following steps:

  1. Read Data File - Supports CSV (.csv) and Excel (.xlsx, .xls) formats
  2. Connect to seekdb - Uses environment variables for server mode, or embedded mode by default
  3. Create Collection - With optional vectorization using default embedding function (all-MiniLM-L6-v2, 384 dimensions)
  4. Import Data - Batch processing with configurable batch size
  5. Verify - Displays record count and data preview after import

User Interaction Guide

For Reading Excel Files

When user wants to preview or inspect an Excel file before importing:

# Preview file structure and data
python scripts/read_excel.py <file_path>

# With column details and statistics
python scripts/read_excel.py <file_path> --columns --stats

This helps users:

  • Understand the file structure (sheets, columns, row count)
  • Identify which column to vectorize
  • Check data quality before importing

For Data Import

When user requests data import, ask:

  1. File path: "Please provide the path to your CSV or Excel file."
    • If user needs sample data, use files from the example-data/ directory
    • Suggest using read_excel.py to preview the file first
  2. Vectorization: "Would you like to enable vector search by vectorizing a column? (yes/no)"
  3. Column selection (if yes): "Which column to vectorize? (e.g., 'Details', 'Description')"
  4. Collection name: "Collection name? (default: derived from filename)"
  5. Connection mode: "Embedded (local) or server mode?"

For Collection Management

  • List collections: Use query_from_seekdb.py list from the querying-from-seekdb skill
  • Delete collection: Run python scripts/import_to_seekdb.py delete <collection_name>

Embedding Functions

The script uses the default embedding function (all-MiniLM-L6-v2, 384 dimensions) when vectorization is enabled via --vectorize-column.

Handling Large Files

For files with >10,000 rows, the import_to_seekdb.py script uses batch processing automatically. You can configure batch size:

python scripts/import_to_seekdb.py import large_file.csv -v Details --batch-size 500

References

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/oceanbase/seekdb-ecology-plugins/importing-to-seekdb">View importing-to-seekdb on skillZs</a>