Brak wyników spełniających kryteria wyszukiwania.

Location: remote with the first day onboarding in Warsaw and occasional visits once per quarter
Rate: up to 180 pln/h on b2b
Project Overview
We are building CaaS (Content as a Service) — a platform that transforms publisher content (PDF textbooks and Excel manifests) into structured, enriched, AI-ready data.
The platform processes content once and exposes it through a unified service layer used by multiple downstream applications.
Key use cases:
RAG-based Teacher Assistant
Editorial tooling
Future AI-powered student-facing products
The goal of this role is to design, build, and maintain a scalable data and AI platform that ingests, processes, enriches, and serves content reliably across multiple environments and consumers.
Responsibilities
Data Engineering & Pipelines
Build and maintain multi-stage data ingestion pipelines
Design and implement idempotent, restartable batch processing workflows
Use S3 as core storage layer for raw and processed data
Implement pipeline stages including:
Content ingestion and book identity assignment
PDF-to-markdown conversion (AI OCR)
Table of contents and structure extraction
Hierarchical chunking
Embedding generation
AI / LLM Processing
Use LLMs and OCR models to extract structured data from PDFs
Design prompts and context strategies for consistent outputs
Generate structured metadata and enrich content for downstream use cases
Data Storage & Consistency
Maintain PostgreSQL (Aurora) as system of record
Design and maintain SQL schemas and versioned migrations
Ensure data consistency across:
S3
PostgreSQL (Aurora)
Vector database (Weaviate)
Implement reconciliation logic across distributed systems
Retrieval & Vector Search
Work with Weaviate for vector search and semantic retrieval
Support RAG-based applications
Design data organization strategies (by subject, country, and client)
APIs & Integration
Build REST APIs using FastAPI
Expose content as a service for multiple downstream applications
Integrate with internal and external systems
Engineering Practices
Write strongly typed Python code (mypy)
Follow CI/CD processes with automated checks (ruff, pytest)
Work across dev / staging / production environments
Debug distributed data inconsistencies
Key Requirements
Must-have
Strong Python development experience (production systems)
AWS experience (S3, Glue, Aurora)
Experience with data pipelines (ETL / batch processing)
Strong SQL and PostgreSQL experience
Experience with schema design and migrations
Experience with LLMs in production (OCR, content processing, enrichment)
Nice-to-have
Prompt engineering / context engineering
Experience with vector databases (Weaviate, Pinecone, Qdrant, pgvector)
Knowledge of embeddings, semantic search, and RAG
Experience with FastAPI
Experience with Airflow / MWAA
Experience building data platforms serving multiple consumers
Zainteresowany ofertą?
Aplikuj już teraz!