
Data Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Brazil.
• Oversee and enhance numerous scrapers that consistently gather court data into the system
• Ingest millions of documents and safely integrate them into the production database
• Assist in the development and advancement of the portal that manages digitization efforts at the United States Library of Congress
• Transform scanned materials into structured text and documents for inclusion in the archive
• Engage in machine learning processes aimed at identifying and eliminating copyrighted content
• Design efficient batch operations for handling five to ten million records
• Troubleshoot slow queries, assess migrations, and offer database support to the team
• Extensive professional experience in Python
• Strong expertise in Django; this framework is essential and the most critical skill
• Solid understanding of database fundamentals, including PostgreSQL, execution plans, query optimization, indexes, and migrations
• Hands-on experience with production web scraping, including the creation, maintenance, and debugging of scrapers
• Familiarity with Celery, Redis, and Kubernetes is a plus
• Experience in applying machine learning to document processing tasks (OCR, classification, extraction) is advantageous
• Background in legal data, civic technology, government technology, or open data is beneficial
• Knowledge of Elasticsearch or similar search infrastructure is a plus
• A genuine interest in law, access to justice, or social justice initiatives is appreciated
• Competitive salary and comprehensive benefits package
• Opportunities for professional development and growth
• Collaborative and inclusive work environment
• Engaging projects that contribute to social impact
Pluribus Digital
GoMining
Get handpicked remote jobs straight to your inbox weekly.