Senior Software Engineer, Short- Message Data Processing

Share on LinkedIn
PythonGO Lang

Long Term - Senior

English - C1 Advanced
Remote

What You’ll Do

  1. Own the build for taking a validated prototype to production — architecture, data model, ingestion design, storage, and integrations.
  2. Build ingestion that survives real-world data: deeply nested export structures (recent single collections have exceeded a million folders), multiple export formats, schema variability, and structural anomalies that must be flagged rather than silently skipped.
  3. Design for volume as a first-class requirement: multi-terabyte collections and tens of millions of records processed without manual splitting, using streaming, chunking, and containerized horizontal scaling.
  4. Model conversation data properly — messages, edits, deletions, replies, reactions, attachments, and time-based participant membership — with stable identifiers and relational links that hold up when a later export rewrites historical records.
  5. Make correctness observable: deterministic quality gates in code, run-level audit trails capturing which build and which settings produced an output, completeness reporting, and explicit reason codes for anything missing. Automated checks are the product, not an afterthought.
  6. Establish engineering foundations: CI/CD, automated testing, observability, coding standards, and security posture on our internal development platform.
  7. Build the self-serve application with the Designer so a processing team — not a single expert — can run jobs, monitor progress, and produce reports in minutes rather than a week.
  8. Work day-to-day with the Product Manager to sequence work, scope complexity, and surface trade-offs, balancing near-term delivery against the platform this becomes. You will be in the room for scoping decisions, not handed a spec.
  9. Absorb and document institutional knowledge from the legacy tool and its author, then retire it via a validated, defensible cutover.
  10. Set the technical bar for engineers who join later as the platform expands to additional data types.


What We’re Looking For

Required

  1. 7+ years of software engineering experience, with substantial time building data-intensive systems in production.
  2. Demonstrated experience designing large-scale data ingestion and processing pipelines — throughput, memory, and parallelism under real constraints; schema variability; deduplication; transformation at scale. You have personally debugged a pipeline that fell over on volume and fixed the design, not just the symptom.
  3. Experience in a domain where provenance and reproducibility are the deliverable — regulated, audited, forensic, financial- reconciliation, clinical, or compliance data. You treat “we can’t explain why these two runs differ” as a defect of the same severity as a crash.
  4. End-to-end ownership of at least one system you built from scratch, including the early architectural decisions that defined it for years.
  5. Strong command of modern application and cloud architecture — you can design and implement a scalable service, including infrastructure, APIs, data stores, and the interface practitioners actually use.
  6. Comfort working closely with product and design. You can hold your own on scope and sequencing, translate technical constraints into terms a PM can decide on, and partner with a designer on an interface rather than bolting one on at the end.
  7. Fluency with modern AI-assisted development. You use these tools daily to move faster, and you have clear judgment about where they belong and where they don’t — our quality gates are deterministic by design.
  8. Startup mindset: comfortable with ambiguity, pragmatic about trade-offs, and able to make decisions that balance speed against long-term soundness.

Strongly Preferred

  1. Experience with messaging, chat, or communications data — Slack, Teams, mobile message extractions, or similar. Conversation threading, participant attribution, and attachment resolution are harder than they look.
  2. Experience replacing a legacy system while it remains in production, including parity validation against the tool being retired.
  3. Background in legal technology, eDiscovery, investigations, or compliance tooling — genuinely useful context, but something a strong engineer can learn here. We are more interested in how you think about completeness than in whether you already know our vocabulary.
  4. Experience building integrations with third-party platforms via API, particularly where you do not control the target system’s
  5. data model.
  6. Experience introducing modern engineering practices (CI/CD, automated testing, observability) in an organization new to formal software development.
  7. Experience working alongside a professional services or consulting delivery team, where your users are practitioners under deadline pressure on high-stakes matters.

Nice to Have

  1. Experience with search and query layers over large datasets (Elasticsearch, Splunk, or similar), including saved and reusable query conditions.
  2. Experience designing modular, plug-in architectures where new data sources can be added without destabilizing existing ones.
  3. Familiarity with entity resolution or identity normalization across inconsistent identifiers — usernames, phone numbers, aliases, email addresses.
  4. Exposure to chain-of-custody and data security requirements in litigation or investigations contexts.
  5. Experience with incremental and delta processing — recomputing only what changed when a dataset is re-delivered or a request expands.
  • 1Tech fit interview
  • 2Interview with final client
  • 3Interview with final client
  • 4Interview with final client
  • 5Interview with final client