Data Scientist

Chantilly, VA
Full Time
VANGUARD
Experienced

Data Scientist

Wyetech is seeking an experienced Data Scientist to support the development, enhancement, and maintenance of enterprise data processing capabilities within a mission-critical Intelligence Community cloud computing environment.

The selected candidate will work closely with Government stakeholders and technical teams to identify requirements for new system capabilities, extend and maintain bulk data pipelines, and enhance multiple applications operating within the customer's cloud infrastructure.

The ideal candidate possesses extensive experience with Python, Spark, PySpark, ETL development, data modeling, SQL databases, and large-scale data processing, along with the ability to transform complex structured and unstructured datasets into reliable, accessible, and actionable information.

This position requires strong analytical capabilities, hands-on data engineering experience, and the ability to collaborate with customers and integration partners in an Agile development environment.

Key Responsibilities

Big Data Processing & Analytics

  • Design, develop, maintain, and enhance large-scale data processing capabilities using Python, Apache Spark, and PySpark.

  • Support the development and optimization of bulk data pipelines within enterprise cloud computing environments.

  • Process, analyze, and transform complex datasets to support mission-critical applications and analytical requirements.

  • Develop scalable data processing solutions capable of handling structured and unstructured information.

  • Perform extensive data reviews and data quality analyses to identify inconsistencies, anomalies, and processing issues.

  • Apply data modeling and transformation techniques to improve data accessibility, consistency, and usability.

  • Support the development of new system capabilities and enhancements to existing cloud-based applications.

  • Collaborate with engineering teams to troubleshoot data processing issues and improve pipeline reliability.

ETL Development & Data Integration

  • Design, develop, implement, and maintain Extract, Transform, Load (ETL) processes supporting enterprise data integration.

  • Perform data mapping, extraction, transformation, and loading across multiple data sources.

  • Integrate disparate structured and unstructured data formats into enriched, query-friendly structured datasets.

  • Develop indexed data files that support efficient querying, reporting, and analytical processing.

  • Process and transform XML, JSON, and other supported data formats.

  • Develop and maintain source-to-target mappings, data dictionaries, and ETL design documentation.

  • Identify and resolve data integration challenges, transformation errors, and data quality issues.

  • Support ongoing maintenance, enhancement, and optimization of bulk data pipelines.

  • Collaborate with integration partners to ensure accurate data ingestion, transformation, and delivery.

Database Development & Data Modeling

  • Develop and execute SQL queries to retrieve, manipulate, validate, and analyze data.

  • Work with relational database technologies, including SQL, MySQL, and PostgreSQL.

  • Support data modeling and analytical development using notebooks and integrated development environments.

  • Utilize Visual Studio and applicable development tools to support data processing and analytical workflows.

  • Develop and maintain structured datasets optimized for analytical queries and reporting.

  • Perform data validation, reconciliation, and quality assurance activities.

  • Identify opportunities to improve database queries, data structures, and processing efficiency.

  • Support the integration of relational database information into enterprise data pipelines.

Log Processing, Monitoring & Data Visualization

  • Process and convert operating system logs and application data logs into actionable reports, metrics, and dashboards.

  • Develop analytical reports and monitoring capabilities using tools such as Amazon CloudWatch and Kibana.

  • Analyze log data to identify patterns, operational trends, anomalies, and performance issues.

  • Apply Regular Expressions (RegEx) to extract, filter, and transform relevant information from complex datasets.

  • Develop metrics and dashboards that improve visibility into system operations and data processing activities.

  • Support reporting solutions that provide actionable insights to technical teams and Government stakeholders.

  • Troubleshoot issues involving log ingestion, data transformation, and reporting accuracy.

  • Maintain and enhance reporting capabilities as customer requirements evolve.

Cloud Data Engineering & Application Support

  • Support data processing and analytical capabilities within enterprise cloud computing environments.

  • Develop and maintain data pipelines supporting multiple cloud-hosted applications.

  • Assist with the integration of cloud services into existing data processing workflows.

  • Support cloud-based data ingestion, transformation, storage, and analytical reporting.

  • Troubleshoot data pipeline failures, processing errors, and application integration issues.

  • Collaborate with cloud engineers and software developers to improve data processing efficiency and system performance.

  • Contribute to the implementation of new cloud-based data capabilities.

  • Support reliable delivery of analytical data to downstream applications and stakeholders.

Agile Development & Customer Collaboration

  • Interface directly with Government customers and integration partners to identify, clarify, and document technical objectives.

  • Translate customer requirements into actionable data processing and analytical development tasks.

  • Participate in Agile development activities, including task definition, scope refinement, planning, and reviews.

  • Collaborate with cross-functional teams to develop and enhance data processing capabilities.

  • Provide technical input regarding data integration requirements, implementation approaches, and development priorities.

  • Document analytical methodologies, data processing workflows, and technical solutions.

  • Support testing, validation, troubleshooting, and continuous improvement of data pipelines and applications.

  • Communicate technical findings, development progress, and potential risks to stakeholders.

Required Qualifications

  • Active TS/SCI security clearance with current Full Scope Polygraph (FSP) is required.

  • Demonstrated experience using Apache Spark, PySpark, and Python for large-scale data processing and analytical development.

  • Demonstrated experience with data mapping, extraction, transformation, and loading (ETL).

  • Experience developing analytical reports using tools such as Amazon CloudWatch and Kibana.

  • Experience processing and converting operating system and application data logs into analytical reports, metrics, and dashboards.

  • Experience using Regular Expressions (RegEx) for data extraction, filtering, and transformation.

  • Experience working with SQL, MySQL, and PostgreSQL.

  • Experience processing and transforming data file formats, including XML and JSON.

  • Experience using integrated development environments and data modeling tools, including notebooks and Visual Studio.

  • Experience transforming disparate structured and unstructured datasets into enriched, query-friendly structured data stored in indexed files.

  • Experience performing extensive data reviews, data validation, and data quality analysis.

  • Experience developing ETL design documentation, including source-to-target mappings and data dictionaries.

  • Experience communicating with customers and integration partners to gather, clarify, and document technical objectives.

  • Experience supporting Agile development activities, including task definition, scope development, and technical reviews.

  • Strong analytical, troubleshooting, and problem-solving skills.

  • Excellent written and verbal communication skills.

  • Ability to work effectively with cross-functional engineering teams and Government stakeholders.

Preferred Qualifications

Advanced Big Data & Analytics

  • Experience deploying analytical capabilities using the Databricks Unified Analytics Platform.

  • Experience with Amazon Elastic MapReduce (EMR) for executing large-scale data processing workloads.

  • Experience joining and processing multiple complex datasets using Apache Spark.

  • Experience tuning Spark streaming and batch processing jobs to improve cluster utilization and processing performance.

  • Experience developing and deploying complex notebook-based data processing pipelines.

  • Experience using Python data analysis libraries, including Pandas.

Cloud Services & DevOps

  • Experience utilizing cloud services such as:

    • AWS Lambda

    • Amazon Simple Notification Service (SNS)

    • Amazon Simple Queue Service (SQS)

    • Amazon Elastic Compute Cloud (EC2)

  • Experience with DevOps and cloud infrastructure tools, including:

    • Amazon CloudWatch

    • AWS Lambda

    • Amazon SQS

    • Amazon DynamoDB

    • Amazon Relational Database Service (RDS)

  • Experience supporting cloud-based data integration, automation, and processing workflows.

  • Familiarity with scalable cloud architectures and distributed data processing environments.

Elasticsearch & Log Analytics

  • Experience working with the Elastic Stack (ELK), including:

    • Elasticsearch

    • Logstash

    • Kibana

  • Experience developing log ingestion, indexing, querying, and analytical reporting capabilities.

  • Experience integrating Elasticsearch with enterprise data processing and monitoring solutions.

  • Experience developing dashboards and visualizations for complex operational datasets.

  • Familiarity with performance monitoring, log analytics, and data-driven operational reporting.

Additional Technical Capabilities

  • Develop and enhance data processing solutions supporting large-scale enterprise applications.

  • Maintain and optimize bulk data pipelines within cloud computing environments.

  • Implement data transformation processes that improve data consistency, accessibility, and usability.

  • Develop reusable analytical workflows and data processing components.

  • Support integration of multiple structured and unstructured data sources.

  • Analyze data quality issues and implement corrective measures.

  • Develop documentation supporting data lineage, source-to-target mappings, and data dictionaries.

  • Support reporting and visualization capabilities for operational and analytical datasets.

  • Collaborate with engineering teams to troubleshoot complex data processing and integration challenges.

  • Support the development and deployment of new data processing capabilities based on evolving customer requirements.

  • Contribute to Agile development activities and continuous improvement initiatives.

Clearance Requirements

Active TS/SCI security clearance with current Full Scope Polygraph (FSP) is required.

Due to federal contract requirements, United States Citizenship and position-appropriate security clearance are required.

The Wyetech Advantage

At Wyetech, you’ll be at the center of an award-winning corporate culture, breaking technological barriers and solving real-world problems for our federal government customers. We are committed to hiring the best of the best, and in return, we offer a world-class, truly unique employee experience that is rare within our industry.

Wyetech believes in generously supporting employees as they prepare for retirement. The company automatically contributes 20% of each employee's gross compensation to a Simplified Employee Pension (SEP) IRA, with no requirement for employee matching. All contributions are fully vested from day one, ensuring immediate ownership of retirement funds.

Additional benefits include:

  • Wyetech provides a generous PTO plan of up to 200 hours annually, aligned with applicable state leave regulations. Employees have the flexibility to adjust their PTO allocation at the start of each calendar year, ensuring it meets their evolving needs.

Full-time employees have the option to participate in a variety of voluntary benefit plans including:

  • A Choice of Medical Plan Options, some with Health Savings Account (HSA)

  • Vision and Dental

  • Life and AD&D Benefits

  • Short and Long-Term Disability

  • Hospital Indemnity, Accident, and Critical Illness Insurances

  • Optional Identity Theft and Legal Protection Services

  • Employee Referral Bonus Eligibility up to $10,000

  • Mobility Among Wyetech-supported Contracts

  • Various team-building events throughout the year such as monthly lunches, summer company picnic, and an annual holiday party.

  • Employees receive two complimentary branded clothing orders annually.

Salary Range: $99.42 to $134.64 per hour

Hourly pay rates listed for this position serve as a general guideline and are not a guarantee of compensation. Compensation will vary dependent upon factors including but not limited to Government contract rates; education; relevant prior work experience, knowledge, skills, and competencies; certifications; and geographic location. Hourly pay rates reflect the pre-benefit gross wage amounts.

Wyetech, LLC is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Affirmative Action Statement:

Wyetech, LLC is committed to the principles of affirmative action in all hiring and employment for minorities, women, individuals with disabilities, and protected veterans.

Accommodations:

Wyetech, LLC is committed to providing an inclusive and accessible hiring process. If you need any accommodations during the application or interview process, please contact Brittney Wood. at 844-WYETECH x727 or staffing@wyetech.com. We are happy to provide reasonable accommodations to ensure equal access to all candidates.

Wyetech Terms and Conditions:

1. Wyetech may reach out to you via SMS messages during the course of your Wyetech candidacy.

2. You can cancel the SMS service at any time. Just text "STOP" to the short code. After you send the SMS message "STOP" to us, we will send you an SMS message to confirm that you have been unsubscribed. After this, you will no longer receive SMS messages from us. If you want to join again, just sign up as you did the first time and we will start sending SMS messages to you again.

3. If you are experiencing issues with the messaging program you can reply with the keyword HELP for more assistance, or you can get help directly at IT@wyetech.com, 844-WYETECH (844-993-8324).

4. Carriers are not liable for delayed or undelivered messages

5. As always, message and data rates may apply for any messages sent to you from us and to us from you. If you have any questions about your text plan or data plan, it is best to contact your wireless provider.

6. If you have any questions regarding privacy, please read our privacy policy below.

Privacy Policy:

We collect your mobile number and related opt-in data for the delivery of SMS messages. Your mobile opt-in data and consent will not be shared with third parties or affiliates for marketing or promotional purposes.

Share

Apply for this position

Required*
We've received your resume. Click here to update it.
Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) or Paste resume

Paste your resume here or Attach resume file

Human Check*