People Data Labs Logo

People Data Labs

Senior Data Acquisition Engineer

Reposted 10 Days Ago
Remote
Hiring Remotely in USA
160K-200K
Senior level
Remote
Hiring Remotely in USA
160K-200K
Senior level
Lead the development of data acquisition and web crawling technologies, ensuring high data quality and optimizing data pipelines for scalability and efficiency.
The summary above was generated by AI

Note for all engineering roles: with the rise of fake applicants and AI-enabled candidate fraud, we have built in additional measures throughout the process to identify such candidates and remove them.

About Us

People Data Labs (PDL) is the provider of people and company data. We do the heavy lifting of data collection and standardization so our customers can focus on building and scaling innovative, compliant data solutions. Our sole focus is on building the best data available by integrating thousands of compliantly sourced datasets into a single, developer-friendly source of truth. Leading companies across the world use PDL’s workforce data to enrich recruiting platforms, power AI models, create custom audiences, and more.

We are looking for individuals who can balance extreme ownership with a “one-team, one-dream” mindset. Our customers are trying to solve complex problems, and we only help them achieve their goals as a team. Our Data Engineering & Acquisition Team ensures our customers have standardized and high quality data to build upon. 

You will be crucial in accelerating our efforts to build standalone data products that enable data teams and independent developers to create innovative solutions at massive scale. In this role, you will be working with a team to continuously improve our existing datasets as well as pursuing new ones. If you are looking to be part of a team discovering the next frontier of data-as-a-service (DaaS) with a high level of autonomy and opportunity for direct contributions, this might be the role for you. We like our engineers to be thoughtful, quirky, and willing to fearlessly try new things. Failure is embraced at PDL as long as we continue to learn and grow from it.

What You Get to Do

  • Use and develop web crawling technologies to capture and catalog data on the internet
  • Support and improve our web crawling infrastructure
  • Structure, define, and model captured data, providing semantic data definition and automate data quality monitoring for data that we crawl
  • Develop new techniques to increase speed, efficiency, scalability, and reliability of web crawls
  • Use big data processing platform to build data pipelines, publish data, and ensure the reliable availability of data that we crawl
  • Work with our data product and engineering team to design and implement new data products with captured data, and enhance and improve upon existing products

The Technical Chops You’ll Need

  • 7+ years industry experience with clear examples of strategic technical problem solving and implementation
  • Strong software development architecture and fundamentals for backend applications
  • Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request/response)
  • Solid programming experience: strong grasp of object-oriented design and experience building applications using asynchronous programming paradigms (e.g., async/await, event loops, or concurrency libraries)
  • Experience building crawlers
  • Proficient in Linux / Unix command line utilities, Linux system administration, architecture, and resource management
  • Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)

People Thrive Here Who Can

  • Must thrive in a fast paced environment and be able to work independently
  • Can work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
  • Strong written communication skills on Slack/Chat and in documents
  • You are experienced in writing data design docs (pipeline design, dataflow, schema design)
  • You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders

Some Nice To Haves

  • Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
  • Experience as a Red Teamer
  • Experience working in data acquisition
  • Experience in network architecture and how to debug and inspect network traffic (DNS, IPv4, Proxies, Application ports and interfaces; packet capture and analysis)
  • Experience with Apache Spark
  • Experience with SQL, including writing advanced queries (e.g., window functions, CTEs)
  • Experience with streaming data platforms (e.g. Kafka or other pub/sub; Spark streaming or other stream processing)
  • Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
  • Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
  • Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
  • Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
  • Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)

Our Benefits

  • Stock
  • Competitive Salaries
  • Unlimited paid time off
  • Medical, dental, & vision insurance 
  • Health, fitness, and office stipends
  • The permanent ability to work wherever and however you want

Comp: $160K - $200K

People Data Labs does not discriminate on the basis of race, sex, color, religion, age, national origin, marital status, disability, veteran status, genetic information, sexual orientation, gender identity or any other reason prohibited by law in provision of employment opportunities and benefits.

Qualified Applicants with arrest or conviction records will be considered for Employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.

Personal Privacy Policy for California Residents
https://www.peopledatalabs.com/pdf/privacy-policy-and-notice.pdf

Top Skills

Spark
Azure)
Big Data Processing Platforms
Cloud Computing Services (Aws
Databricks
GCP
Linux
SQL
Web Crawling Technologies

Similar Jobs

48 Minutes Ago
Remote or Hybrid
Kirkland, WA, USA
164K-286K Annually
Mid level
164K-286K Annually
Mid level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Manage and mentor a team of ServiceNow platform developers, oversee project management, conduct architectural reviews, and ensure system optimization and availability.
Top Skills: JavaJavaScriptLinuxPythonServicenow
48 Minutes Ago
Remote or Hybrid
San Diego, CA, USA
147K-258K Annually
Senior level
147K-258K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
The Staff Software Engineer will build scalable applications, collaborate with teams, and enhance features while ensuring high-quality software development in an Agile environment.
Top Skills: AngularjsApi GatewaysAWSAzureCatchpointCdnsCSSDell BoomiHTMLJavaScriptKafkaReactSplunkSQL
48 Minutes Ago
Remote or Hybrid
Orlando, FL, USA
2-2
Junior
2-2
Junior
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
The Associate Customer Success Manager acts as a customer advocate, overseeing a customer portfolio to ensure product adoption and resolve escalations using relevant company resources.
Top Skills: Ai-Enhanced TechnologyAi-Powered Tools

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account