Databricks Databricks-Certified-Data-Engineer-Professional Prüfungsthemen:
| Abschnitt | Gewichtung | Ziele |
|---|---|---|
| Thema 1: Datenverwaltung und Governance | 7% | - Verwaltung von Datenobjekten und Metadaten - Umsetzung von Richtlinien und Standards im Datenbereich - Verwendung von Unity Catalog für die Datenverwaltung |
| Thema 2: Datentransformation, Bereinigung und Qualitätssicherung | 10% | - Einhaltung von Standards zur Datenqualität - Anwenden von Regeln zur Datenbereinigung und -validierung - Umsetzung von Schemaentwicklung und -verwaltung |
| Thema 3: Kosten- und Leistungsoptimierung | 13% | - Anwendung bewährter Verfahren zur Kostensteuerung - Verbesserung der Leistung von Abfragen und Pipelines - Optimierung von Rechen- und Speicherressourcen |
| Thema 4: Datenmodellierung | 6% | - Gestaltung der Medallion-Architektur - Umsetzung dimensionaler und relationaler Modelle - Optimierung des Tabellenaufbaus und der Partitionierung |
| Thema 5: Entwicklung von Code zur Datenverarbeitung mit Python und SQL | 22% | - Verwendung von Databricks-spezifischen Bibliotheken und APIs - Umsetzung komplexer Datenverarbeitungslogiken - Schreiben von effizientem und wartbarem Code |
| Thema 6: Datenaufnahme und -erfassung | 7% | - Aufnahme von Daten aus unterschiedlichen Quellen - Verarbeitung von inkrementellen und stapelweisen Datenladungen - Verwendung von Auto Loader und strukturiertem Streaming |
| Thema 7: Überwachung und Warnmeldungen | 10% | - Einrichtung von Warnmeldungen und Benachrichtigungen - Nachverfolgung von Datenherkunft und Messwerten - Überwachung der Leistung und Betriebsfähigkeit von Pipelines |
| Thema 8: Fehleranalyse und Bereitstellung | 10% | - Fehlerbehebung und Analyse von Pipelines - Bereitstellung mithilfe von Asset Bundles, CLI und APIs - Anwendung von CI/CD- und DevOps-Verfahren |
| Thema 9: Datenfreigabe und föderierter Datenzugriff | 5% | - Sichere Datenfreigabe mithilfe von Delta Sharing - Verwaltung des plattformübergreifenden Datenzugriffs - Umsetzung von Lakehouse Federation |
| Thema 10: Gewährleistung von Datensicherheit und Einhaltung von Vorschriften | 10% | - Umsetzung von Zugriffskontrollen und Berechtigungen - Gewährleistung des Datenschutzes und der Einhaltung rechtlicher Vorgaben - Sicherung ruhender und übertragener Daten |
Databricks Certified Data Engineer Professional Databricks-Certified-Data-Engineer-Professional Prüfungsfragen mit Lösungen
To identify the top users consuming compute resources, a data engineering team needs to monitor usage within their Databricks workspace for better resource utilization and cost control.
The team decided to use Databricks system tables, available under the System catalog in Unity Catalog, to gain detailed visibility into workspace activity. Which SQL query should the team run from the System catalog to achieve this?
- A. SELECT sku_name,
usage_metadata.run_name AS user_email,
SUM(usage_quantity) AS total_dbus
FROM system.billing.usage
GROUP BY user_email, sku_name
ORDER BY total_dbus DESC
LIMIT 10 - B. SELECT sku_name,
identity_metadata.created_by AS user_email,
COUNT(usage_quantity) AS total_dbus
FROM system.billing.usage
GROUP BY user_email, sku_name
ORDER BY total_dbus DESC
LIMIT 10 - C. SELECT sku_name,
identity_metadata.created_by AS user_email,
SUM(usage_quantity * usage_unit) AS total_dbus
FROM system.billing.usage
GROUP BY user_email, sku_name
ORDER BY total_dbus DESC
LIMIT 10 - D. SELECT identity_metadata.run_as AS user_email,
SUM(usage_quantity) AS total_dbus
FROM system.billing.usage
GROUP BY user_email
ORDER BY total_dbus DESC
LIMIT 10
Erklärung: (Nur für DeutschPrüfung-Mitglieder sichtbar)
Which of the following technologies can be used to identify key areas of text when parsing Spark Driver log4j output?
- A. pyspsark.ml.feature
- B. Scala Datasets
- C. Regex
- D. C++
- E. Julia
Erklärung: (Nur für DeutschPrüfung-Mitglieder sichtbar)
What statement is true regarding the retention of job run history?
- A. It is retained for 90 days or until the run-id is re-used through custom run configuration
- B. It is retained until you export or delete job run logs
- C. It is retained for 30 days, during which time you can deliver job run logs to DBFS or S3
- D. It is retained for 60 days, during which you can export notebook run results to HTML
- E. It is retained for 60 days, after which logs are archived
A Delta Lake table representing metadata about content posts from users has the following schema:
user_id LONG, post_text STRING, post_id STRING, longitude FLOAT,
latitude FLOAT, post_time TIMESTAMP, date DATE
This table is partitioned by the date column. A query is run with the following filter:
longitude < 20 & longitude > -20
Which statement describes how data will be filtered?
- A. Statistics in the Delta Log will be used to identify data files that might include records in the filtered range.
- B. The Delta Engine will scan the parquet file footers to identify each row that meets the filter criteria.
- C. No file skipping will occur because the optimizer does not know the relationship between the partition column and the longitude.
- D. Statistics in the Delta Log will be used to identify partitions that might Include files in the filtered range.
- E. The Delta Engine will use row-level statistics in the transaction log to identify the flies that meet the filter criteria.
Erklärung: (Nur für DeutschPrüfung-Mitglieder sichtbar)
A data engineer is optimizing a MERGE operation on an 800GB UC-managed table that experiences frequent updates and deletions. Which two actions should the engineer prioritize to improve MERGE performance? (Choose two.)
- A. Partition the table by date.
- B. Overwrite the table instead of Merge.
- C. Use ZORDER on high-cardinality columns.
- D. Apply liquid clustering using the merge join keys.
- E. Enable deletion vectors on the table if not already enabled.
Erklärung: (Nur für DeutschPrüfung-Mitglieder sichtbar)






1182 Kundenbewertungen

