J

Big Data Application Developer

Jobgether • United State
Remote
Apply
AI Summary

Design and develop large-scale data solutions supporting advanced analytics, reporting, and machine learning initiatives. Build and optimize high-volume data pipelines across modern big data ecosystems ensuring reliability, scalability, and performance. Requires 5+ years of experience with Apache Spark, Hadoop ecosystem, and distributed data platforms.

Key Highlights
Design and develop scalable data applications and processing pipelines
Strong expertise with Apache Spark, Hadoop ecosystem (Hive, HDFS, Sqoop, HBase), and streaming platforms (Kafka, Spark Streaming, Flink)
Proficiency in Python or Shell scripting, SQL, and workflow orchestration tools (Airflow, Oozie)
Key Responsibilities
Design, develop, and maintain scalable data applications and processing pipelines
Build and maintain data ingestion, transformation, and analytics workflows supporting enterprise applications
Develop high-performance processing solutions using Apache Spark with Scala, Python, or Java
Work with Hadoop technologies including Hive, HDFS, Sqoop, HBase, and related distributed data platforms
Implement and support streaming data solutions using technologies such as Kafka, Spark Streaming, or Flink
Optimize data processing performance through effective use of distributed computing concepts
Develop and maintain workflow automation using orchestration tools such as Airflow or Oozie
Create scripts and automation tools using Python or Shell to improve operational efficiency
Troubleshoot complex data platform issues, perform debugging, and maintain technical documentation
Support cloud-based big data environments and contribute to modernization initiatives
Technical Skills Required
Apache Spark Hadoop ecosystem Python
Benefits & Perks
Competitive annual salary range of $100,000-$150,000
Fully remote work opportunity within the United States
Full-time direct employment
Nice to Have
Preferred experience with cloud-based Hadoop platforms such as AWS EMR, Azure HDInsight, or Databricks
Familiarity with modern lakehouse technologies including Delta Lake, Iceberg, or Hudi
Exposure to data governance tools such as Apache Atlas or Collibra
Experience with Kubernetes-based data platforms, CI/CD practices, and infrastructure-as-code solutions

Job Description


This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Big Data Application Developer based in United States.

This role offers the opportunity to design and develop large-scale data solutions that support advanced analytics, reporting, and machine learning initiatives.

You will build and optimize high-volume data pipelines across modern big data ecosystems while ensuring reliability, scalability, and performance.

The position combines software engineering expertise with deep knowledge of distributed data platforms and processing frameworks.

You will work on complex data challenges involving structured and unstructured information, helping organizations transform data into actionable insights.

The ideal candidate will contribute to enterprise-level data platforms while collaborating with technical teams to deliver efficient and innovative solutions.

This is a strong opportunity for a data engineering professional looking to work with cutting-edge technologies in a remote environment.

Accountabilities

The Big Data Application Developer will be responsible for designing, developing, and maintaining scalable data applications and processing pipelines. This role requires strong technical expertise across big data technologies, distributed systems, and data engineering practices to support reliable analytics and business intelligence solutions.

  • Design, develop, and operate large-scale data processing pipelines within Hadoop-based ecosystems.
  • Build and maintain data ingestion, transformation, and analytics workflows supporting enterprise applications.
  • Develop high-performance processing solutions using Apache Spark with Scala, Python, or Java.
  • Work with Hadoop technologies including Hive, HDFS, Sqoop, HBase, and related distributed data platforms.
  • Implement and support streaming data solutions using technologies such as Kafka, Spark Streaming, or Flink.
  • Optimize data processing performance through effective use of distributed computing concepts, including partitioning, replication, and fault tolerance.
  • Develop and maintain workflow automation using orchestration tools such as Airflow or Oozie.
  • Write efficient SQL queries and work with relational and NoSQL databases to support data-driven applications.
  • Create scripts and automation tools using Python or Shell to improve operational efficiency.
  • Troubleshoot complex data platform issues, perform debugging, and maintain technical documentation.
  • Support cloud-based big data environments and contribute to modernization initiatives when required.

Requirements

The ideal candidate brings extensive experience building and supporting enterprise-scale data platforms, with strong knowledge of big data technologies and software engineering principles. This professional should have the ability to solve complex technical problems while delivering reliable and scalable data solutions.

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of professional experience designing and operating big data pipelines.
  • Strong hands-on experience with Apache Spark in production environments using Scala, Python, or Java.
  • Solid expertise with Hadoop ecosystem technologies, including Hive, HDFS, Sqoop, and HBase.
  • Experience with streaming platforms such as Kafka, Spark Streaming, or Flink.
  • Strong SQL skills and experience working with relational and NoSQL databases.
  • Experience with workflow orchestration tools such as Airflow or Oozie.
  • Strong understanding of distributed systems architecture and data processing principles.
  • Proficiency in Python or Shell scripting.
  • Excellent troubleshooting, debugging, analytical, and documentation skills.
  • Preferred experience with cloud-based Hadoop platforms such as AWS EMR, Azure HDInsight, or Databricks.
  • Familiarity with modern lakehouse technologies including Delta Lake, Iceberg, or Hudi.
  • Exposure to data governance tools such as Apache Atlas or Collibra.
  • Experience with Kubernetes-based data platforms, CI/CD practices, and infrastructure-as-code solutions.

Benefits

  • Competitive annual salary range of $100,000-$150,000, depending on experience and qualifications.
  • Fully remote work opportunity within the United States.
  • Full-time direct employment opportunity.
  • Opportunity to work on large-scale data engineering and analytics initiatives.
  • Exposure to modern big data technologies and enterprise data platforms.
  • Career growth opportunities within a technology-focused environment.
  • Opportunity to collaborate with experienced engineers and technical teams.
  • Inclusive workplace committed to equal employment opportunities.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.


Similar Jobs

Explore other opportunities that match your interests

AI Business Analyst - Healthcare Data

Data Science
•
2h ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

wholesum billing

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

machinify

United State
Visa Sponsorship Relocation Remote
Job Type Contract
Experience Level Mid-Senior level

Revel IT

United State

Subscribe our newsletter

New Things Will Always Update Regularly